Summary
- Slurm runs over 60% of TOP500 supercomputers, but stock Slurm-on-Kubernetes uses partitions, accounts, and namespaces that are not security boundaries for paying tenants.
- Commercial HPC isolation requires control plane isolation (dedicated scheduler/API/state) plus data plane isolation (dedicated Private Nodes) — neither alone is sufficient.
- The production architecture uses four layers: a per-tenant control plane, Private Nodes, kernel-native workload hardening, and Netris-backed network isolation; Lintasarta reached 170+ tenant clusters in production within 90 days.
- To deploy this on your GPU estate without building the stack yourself, vCluster Platform provisions isolated Slurm tenant clusters with dedicated Private Nodes via GitOps.
Every major GPU cloud is converging on the same architecture: Slurm, the scheduler running over 60% of TOP500 supercomputers, running on a Kubernetes substrate. SchedMD shipped Slinky, the Kubernetes-native Slurm operator, as the upstream answer.
The trend is correct. The isolation problem is not solved by following it.
A stock Slurm-on-Kubernetes deployment does not give each paying customer their own isolated environment. Slurm partitions and accounts are accounting constructs. Kubernetes namespaces are policy-level separations. Neither is a security boundary for untrusted, commercial tenants running GPU workloads.
This article is about the engineering required to close that gap: how to give each HPC customer their own scheduler, their own queue, their own node set, and a hard hardware boundary, without provisioning a separate physical cluster per customer.
Why Standard Slurm-on-Kubernetes Fails for Commercial Tenants
Slurm was designed for a trusted internal environment: one slurmctld controlling one homogeneous node pool for a single organization's users. The multi-cluster and partition features that Slurm provides are resource allocation mechanisms for trusted users at one site. They were never designed to enforce security isolation between untrusted, paying customers.
When you deploy Slurm on Kubernetes for multiple commercial tenants, you inherit the same problem at two layers simultaneously.
At the Slurm layer, a single slurmctld manages scheduler state for every tenant. One tenant's jobs are visible to the scheduler managing another tenant's queue. Slurm's accounting database (slurmdbd) holds cross-tenant data. Partitions define which nodes a user can submit to; they do not prevent a user from observing cluster-wide state.
At the Kubernetes layer, the gaps are structural. Namespaces and RBAC are policy-level separations inside a control plane, not a hardware boundary. They do not provide the per-tenant API, scheduler, or dedicated node set that commercial GPU services require. The production requirement is per-tenant control planes and dedicated Private Nodes.
GPU contention is the most immediate operational consequence. When tenant workloads are not pinned to dedicated Private Nodes, they contend for NUMA topology, PCIe bandwidth, and memory pressure. An AI training job that depends on consistent GPU throughput has no performance guarantee without dedicated Private Nodes.
For regulated workloads (sovereign AI, financial services HPC, healthcare compute), control planes and nodes that are not dedicated per tenant do not satisfy the requirement. For any commercial AI cloud serving paying customers, that gap creates an unacceptable liability.
What Real Tenant Isolation Requires
The requirement for commercial Slurm-on-Kubernetes breaks into two distinct layers: control plane isolation and data plane isolation. Both must be enforced. Neither alone is sufficient.
Control plane isolation means each tenant has an independent scheduler, independent state, and independent API surface. No tenant can observe another tenant's queue, job history, or node inventory. The control plane itself is invisible to tenants; they cannot access etcd, cannot read control plane logs, and cannot see the infrastructure underneath their environment.
Data plane isolation means each tenant's workloads run on dedicated physical machines, isolated at the hardware boundary from every other tenant. This is the only model that eliminates noisy-neighbor GPU contention and contains the blast radius of a kernel-level vulnerability or a container escape.
These two requirements together define what a genuine tenant cluster is for commercial HPC. Meeting both is the engineering problem. The sections below describe a four-layer model that addresses them.
The Four-Layer Isolation Model for Slurm Tenant Clusters
Layer 1: A Dedicated Control Plane Per Tenant
The foundation is control plane virtualization. vCluster gives each tenant their own CNCF-certified Kubernetes control plane: a dedicated API server, etcd instance, scheduler, controllers, and RBAC scope, running independently from every other tenant. The control plane cluster is completely invisible to the tenant. They cannot see its logs, cannot reach its etcd, and have no visibility into the infrastructure layer underneath.
Inside this isolated Kubernetes control plane, the Slinky slurm-operator is deployed. Slinky is SchedMD's Kubernetes-native Slurm operator, and it models each Slurm daemon (slurmctld, slurmdbd, slurmd, slurmrestd) as a Kubernetes Custom Resource and a pod. Because these daemons run inside vCluster's isolated control plane, the tenant's Slurm scheduler has no visibility into any other tenant's Slurm environment.
The tenant's Slurm control plane (their scheduler, their accounting database, their REST API endpoint) is entirely their own.
Layer 2: Private Nodes for Hardware-Level Data Plane Isolation
Control plane isolation solves the scheduler boundary. It does not solve the node boundary. For that, the production default is Private Nodes.
With Private Nodes, dedicated worker nodes are joined directly into the tenant's isolated control plane. Worker nodes are not drawn from the control plane cluster. The tenant's workloads run on separate physical machines that belong entirely to their environment. This delivers hardware-level isolation approaching a dedicated physical cluster, without the provisioning cost or time of one.
The operational consequences are concrete:
- Noisy-neighbor GPU contention is eliminated. The tenant's GPUs are not shared with any other tenant's processes at any layer.
- Each tenant gets their own CNI. One tenant can run SR-IOV with Multus for GPU-direct RDMA traffic; another can run Calico with eBPF. Dedicated Private Nodes make simultaneous per-tenant CNI configurations possible.
- Each tenant's storage drivers, kernel parameters, and node-level tooling are scoped to their nodes. The platform operator can deploy daemonsets or operators on those nodes that are invisible to the tenant, enforced through admission policies rather than RBAC (which is purely additive and has no deny rules).
Private Nodes must be enabled at cluster creation via privateNodes.enabled: true, as documented in the Private Nodes guide. Every tenant cluster runs on dedicated Private Nodes from the start.
Layer 3: vNode for Kernel-Native Workload Hardening
Dedicated nodes establish the hardware boundary. They do not, by themselves, isolate the processes running within a tenant's workloads from one another. For commercial HPC (where a tenant may run code from multiple end users, install packages dynamically, or require root access inside a job), an additional isolation layer at the workload level is required.
vNode provides this using seccomp, cgroups, and Linux namespaces. It is purpose-built for safely running untrusted code, including workloads that require dynamic package installs or elevated privileges inside the container. There is no guest kernel, no hypervisor, and no virtual kubelet. Bare-metal GPU performance is preserved.
vNode isolates the workload process, not the GPU silicon. GPU partitioning (MIG, vGPU, time-slicing, or DRA) operates beneath this layer, within the tenant's Private Node pool, and is managed by the NVIDIA GPU Operator and the NVIDIA DRA driver. These are orthogonal capabilities: GPU partitioning subdivides a physical GPU within a tenant's node set; vNode hardens the process executing against that GPU.
Layer 4: Network and Hardware Isolation via vMetal and Netris
The isolation envelope extends to the network fabric. vMetal is the Infrastructure Orchestrator: the machine layer with one stable, EC2-like API that turns raw racks into Bare Metal Machines and Virtual Machines, handling PXE boot, OS installation, machine registration, and full node lifecycle management. vMetal integrates Metal3, KubeVirt, NVIDIA NICo, VMware, OpenStack, and Netris, and supports zero-touch provisioning.
Network isolation is delivered by Netris, which vMetal orchestrates to create per-tenant VLANs, VRFs, and ACLs. DPU policies enforce traffic segregation at the NIC before it reaches the fabric. Each tenant's network plane is programmatically scoped and does not overlap with any other tenant's. This is Netris's role in the stack: network isolation is not a native vCluster or vNode capability.
The result of all four layers combined: each tenant has their own Slurm scheduler, their own job queue, their own dedicated GPUs on their own nodes, their own network segment, and a workload isolation boundary around every process their jobs execute.
What This Looks Like in Practice
Provisioning a Tenant Slurm Cluster
Platform operators define a reusable VirtualClusterTemplate (a version-controlled CRD that captures the standard configuration for a tenant cluster). The template encodes Private Nodes enabled, resource quotas, CNI selection, and which certified stack to deploy. Templates define what the cluster looks like; Apps define what is pre-installed in it. These are distinct concepts in vCluster Platform.
To provision a new Slurm tenant, an operator declares a VirtualClusterInstance in a Git repository referencing that template. A GitOps controller (Argo CD or Flux) applies the manifest. vCluster Platform provisions the isolated control plane and attaches the dedicated Private Nodes. Argo CD ApplicationSets then automatically install the Slinky slurm-operator, the NVIDIA GPU Operator, and any other tenant-specific tooling into the new cluster.
The provisioning sequence is designed to run with minimal manual intervention, as demonstrated in deployments like Lintasarta's.
The Tenant's View
The tenant receives a kubeconfig pointing to their private, isolated Kubernetes API server. When they run kubectl get nodes, they see only their dedicated node pool. The control plane cluster is not visible. Other tenants' nodes, pods, and workloads do not appear in any query.
Their Slurm environment (slurmctld, slurmdbd, slurmrestd) runs inside their control plane as Slinky-managed pods. They submit jobs with srun and sbatch against their own scheduler. They manage their own queue, their own partitions, their own accounting. From their perspective, they are operating a dedicated HPC cluster. They are not. They are operating a fully isolated tenant cluster on dedicated Private Nodes within the platform's bare metal estate.
Scope Discipline: Attributing Capabilities Correctly
Slurm tenant isolation on Kubernetes is a multi-component stack. Attributing each capability to the correct component is what makes the architecture auditable and maintainable.
SchedMD / Slinky owns Slurm scheduling. The slurm-operator manages the full lifecycle of each tenant's Slurm control plane: daemon deployment, job queue management, partition configuration, and accountancy. Slinky is what makes Slurm Kubernetes-native.
vCluster Platform provides the isolation envelope: control plane virtualization and Private Nodes for hardware-level data plane isolation. vNode provides kernel-native workload hardening. vCluster Platform does not perform Slurm scheduling, gang scheduling, or job queue management. The Slurm cluster runs inside the isolation envelope vCluster Platform provides.
Netris provides network isolation: VLANs, VRFs, ACLs, and DPU policies. This is not a vCluster or vNode capability. vMetal orchestrates Netris as the network isolation fabric.
NVIDIA GPU Operator and DRA manage the GPU enablement stack within each tenant's Private Node pool: driver installation, device plugin deployment, DCGM telemetry, and advanced partitioning via MIG, vGPU, time-slicing, or DRA. GPU partitioning is orthogonal to tenant isolation and layers beneath it.
No single component in this stack delivers the full requirement. The architecture works because each component does exactly what it is designed to do, in its correct position in the stack.
Deployed at Scale
This architecture is in production at AI clouds operating at significant scale.
Lintasarta launched Indonesia's leading GPU cloud on vCluster and reached 170+ tenant clusters in production within 90 days. The speed of that deployment reflects what a GitOps-driven, template-based provisioning model enables at the operations layer.
The isolation model described in this article is validated in NVIDIA's DGX reference architecture for sovereign AI deployments, which vCluster Labs authored.
What to Build Next
The four-layer model (dedicated control plane, Private Nodes, vNode workload hardening, and Netris-backed network isolation) is the architecture that makes Slurm tenant isolation commercially viable on a pooled GPU estate.
Building this without a proven platform is the kind of infrastructure engineering that takes significant time and expertise. The dependency chain is long: control plane virtualization, Private Nodes provisioning, Slinky integration, NVIDIA GPU Operator certified stacks, Netris network configuration, GitOps templating. DIY implementations that replicate this stack remain the domain of the largest, best-resourced engineering teams; even those teams are still building what this platform delivers today.
To see how vCluster Platform provisions isolated Slurm tenant clusters on your GPU infrastructure, schedule an enterprise demo.
Frequently Asked Questions
What is a tenant cluster for Slurm on Kubernetes?
A tenant cluster is a dedicated, isolated environment that gives each customer their own Slurm scheduler, job queue, node set, and Kubernetes control plane. In this architecture, each tenant cluster runs on its own Private Nodes and is provisioned via vCluster Platform, delivering hardware-level isolation without requiring a separate physical cluster per customer.
How does Private Nodes provide hardware-level data plane isolation?
Private Nodes attach dedicated physical worker nodes directly into a tenant’s isolated control plane cluster, ensuring dedicated physical infrastructure for each tenant. This eliminates noisy-neighbor GPU contention, enables tenant-specific CNI and storage configurations, and creates a hard boundary that approximates a dedicated physical cluster.
Why is a dedicated control plane per tenant necessary for commercial AI clouds?
A dedicated control plane cluster prevents any tenant from observing or affecting another tenant’s scheduler state, job queues, or Kubernetes objects. For commercial GPU clouds serving paying customers, this control plane isolation is required to meet security, compliance, and performance expectations; namespaces and RBAC alone are not sufficient for untrusted tenants.
What is the difference between control plane isolation and data plane isolation?
Control plane isolation separates the Kubernetes and Slurm scheduling and state layers per tenant, while data plane isolation separates the physical compute resources that run tenant workloads. Both are required: control plane isolation prevents cross-tenant visibility and API interference, and data plane isolation prevents cross-tenant resource contention and security breaches.
How does vCluster enable tenant isolation for Slurm environments?
vCluster provides each tenant with their own CNCF-certified Kubernetes control plane cluster, including a dedicated API server, etcd, and scheduler. Combined with Private Nodes and vNode workload hardening, vCluster delivers the isolation envelope in which Slurm daemons (managed by Slinky) run; it does not perform Slurm scheduling itself.
Does vNode replace GPU partitioning like MIG or vGPU?
No. vNode hardens the process boundary around workloads using seccomp, cgroups, and Linux namespaces, isolating processes from each other. GPU partitioning (MIG, vGPU, time-slicing, DRA) operates beneath vNode and subdivides a physical GPU within a tenant’s Private Node pool; they are orthogonal capabilities.
How do tenants access their isolated Slurm environment?
Tenants receive a kubeconfig that points exclusively to their private, isolated Kubernetes API server. When they run kubectl get nodes or submit Slurm jobs, they only see their own dedicated nodes and scheduler; the control plane cluster and other tenants’ resources are completely invisible.
Can this architecture be implemented without vCluster Platform?
Yes, but it requires building and integrating control plane virtualization, Private Nodes provisioning, Slinky, NVIDIA GPU Operator, Netris network isolation, and GitOps templating yourself. Few engineering teams have the resources to replicate this stack; the architecture is proven in production AI clouds using vCluster Platform.
Deploy your first virtual cluster today.