Summary
- The gap: In GPU-accelerated AI infrastructure, the most dangerous tenant isolation gaps sit at the API, compute, and network layers, not in the database.
- The production default: Dedicated physical worker nodes per tenant is the production default, so the API, compute, and network blast radius are contained per tenant.
- The architecture: Effective isolation composes separate API, compute, and network boundaries into a single control boundary.
- Scale proof: The full-stack approach is proven at 100K+ GPUs across 50+ GPU Clouds & Fortune 500s, with 40M+ tenant clusters created and production launches in 45-90 days.
- Action: For commercial GPU cloud operators, vCluster Platform delivers per-tenant API servers, etcd, RBAC, CRDs, and dedicated Private Nodes as the production default.
Ninety percent of tenant isolation content stops at the database. Row-level security, connection pooling, and schema separation are real controls, and they matter. But in GPU-accelerated AI infrastructure, the most dangerous isolation gaps live elsewhere: at the API gateway, inside the compute layer, and across the network fabric.
The reason is structural. AI workloads execute untrusted, often model-generated code. In production, Private Nodes are the default: tenants run on dedicated hardware, not shared GPU nodes. A container escape on a shared node is a realistic event with a blast radius that row-level security cannot contain. A flat cluster network means a compromised pod can scan and reach every other tenant's services. A shared API server means a runaway controller or misconfigured CRD from one tenant can destabilize the entire platform.
Shared GPU infrastructure at scale creates a pattern where one tenant's behavior can create a platform-wide incident that floods support and degrades service for everyone else. The API, compute, and network layers were never hardened to the same standard.
This article takes each layer in turn: the threat model, the architecture, and the control that closes the gap.
Layer 1: API Layer (Per-Tenant Control Planes)
The Threat Model
In a standard shared Kubernetes cluster, namespace isolation creates the appearance of separation. In practice, the API server, etcd, and scheduler are shared by every tenant on the platform.
This creates three concrete risks:
- Cross-tenant visibility. A tenant with misconfigured RBAC can query workloads, nodes, and service endpoints belonging to other tenants via the shared API server.
- Control plane instability. A misconfigured CRD or a runaway admission webhook from one tenant can overwhelm etcd or crash the API server for all tenants simultaneously.
- RBAC complexity at scale. Managing RBAC for hundreds of tenants against a single API server is error-prone. A single misconfiguration can grant cluster-scoped privileges.
The Architecture
// Weak isolation: Shared control plane
// Single point of failure, shared blast radius
Kubernetes Control Plane (shared) {
APIServer -> visible to all tenants
etcd -> shared state for all tenants
Scheduler -> shared by all tenants
Namespace: tenant-a { workloads, CRDs ... }
Namespace: tenant-b { workloads, CRDs ... }
}
// Strong isolation: Per-tenant control planes with Private Nodes
// Each tenant is its own independent failure domain
TenantCluster-A {
ControlPlane (dedicated to tenant-a) {
APIServer -> scoped to tenant-a only
etcd -> tenant-a state only
Scheduler -> tenant-a workloads only
RBAC -> tenant-a policies only
}
PrivateNodes (dedicated to tenant-a)
}
TenantCluster-B {
ControlPlane (dedicated to tenant-b) {
APIServer -> scoped to tenant-b only
etcd -> tenant-b state only
Scheduler -> tenant-b workloads only
RBAC -> tenant-b policies only
}
PrivateNodes (dedicated to tenant-b)
}The Control
vCluster Platform provides each tenant with a fully isolated control plane: their own API server, etcd, scheduler, controllers, RBAC, admission control, and namespace-scoped CRDs, with dedicated Private Nodes joined to each tenant cluster. A misconfigured CRD, a runaway controller, or a compromised API credential in one tenant has zero blast radius on any other.
The underlying control plane cluster is completely invisible to the tenant. They interact with their own dedicated API endpoint as if it were a private physical cluster (the GKE model, where the control plane is not exposed).
This model extends beyond Kubernetes. vCluster Platform virtualizes the control plane for any cluster type: Slurm, NVIDIA Run:ai, and Ray run as isolated tenant clusters with the same control plane separation; Inference and Agent Sandbox will extend the same model.
Layer 2: Compute Layer (Hardware Isolation and Container Breakout Prevention)
The Threat Model
The compute layer is where AI infrastructure diverges most sharply from conventional SaaS. AI agents and inference workloads regularly execute dynamic code, install packages, and run with elevated privileges. On a shared node, that execution environment is one kernel vulnerability away from a full host compromise.
Three specific risks dominate:
- Container escape. A workload exploiting a kernel misconfiguration can break out of its container and access the host node, reading other tenants' process memory, filesystem data, or GPU state.
- Privilege escalation. Monitoring agents and GPU drivers frequently require privileged containers. If compromised, a privileged container on a shared node provides direct access to host-level resources.
- Noisy-neighbor resource exhaustion. Without strict per-tenant resource controls, one tenant's GPU-intensive workload can monopolize compute, memory, and I/O bandwidth, degrading performance for every other tenant on the node. This is a well-understood operational failure mode: latency spikes, queue buildup, and support escalations that affect the entire platform.
The Architecture
// Hardware isolation: Private Nodes (production default)
// Each tenant gets dedicated physical nodes: no shared kernel, no shared compute
TenantCluster-A {
PrivateNode-1 (dedicated to tenant-a) {
Pod-A1, Pod-A2
// No tenant-b workloads ever scheduled here
}
}
TenantCluster-B {
PrivateNode-2 (dedicated to tenant-b) {
Pod-B1, Pod-B2
}
}
// Kernel-native isolation: vNode (defense-in-depth layer)
// Sandboxes workloads without hypervisor overhead
Host Node {
vNode-Sandbox-A (tenant-a) {
// Separate user, pid, net, mount namespaces
// seccomp filter applied per-process
Container-A1, Container-A2
// A container escape is contained within this sandbox
// Cannot reach the host or vNode-Sandbox-B
}
vNode-Sandbox-B (tenant-b) {
Container-B1, Container-B2
}
}The Controls
Private Nodes: hardware-level isolation as the production default
vCluster Platform's production default is Private Nodes: dedicated physical worker nodes joined directly and privately into each tenant cluster. No workloads from other tenants ever share the same hardware, kernel, or OS. This delivers dedicated-hardware isolation without requiring entirely separate physical clusters, and it is what vCluster Platform deploys by default for enterprise and commercial GPU cloud use cases.
Private Nodes are set at tenant cluster creation. The dedicated nodes are not visible from the control plane cluster; they exist only within that tenant cluster. (Where cross-network connectivity is needed, the built-in vCluster VPN feature can join them over encrypted tunnels.) With Private Nodes, each tenant cluster also gets its own CNI and storage stack, which carries direct implications for the network layer covered below.
vNode (optional defense-in-depth layer)
For defense-in-depth on top of Private Nodes, vNode adds per-process isolation using seccomp, cgroups, and Linux namespaces.
vNode achieves this through three kernel-level mechanisms: Linux user namespaces, cgroups, and targeted seccomp filtering. Together they ensure that a process inside a vNode sandbox cannot access another tenant's files, processes, or hardware information. This protection holds even after a container escape. A breakout lands inside the vNode sandbox, not on the host.
Because there is no hypervisor and no guest kernel, GPU workloads run at bare metal speed. The NVIDIA GPU Operator handles GPU device injection through standard Kubernetes device plugin mechanisms, and vNode is compatible with it via RuntimeClass and CDI annotations, so that injection works correctly inside the sandboxed environment. vNode's role is sandboxing the workload process; it does not manage GPU device assignment. vNode is hardened and pentested to the kernel.
One operational note: when using newer NVIDIA GPU Operator versions on shared nodes with vNode, containerd 2.x is required for correct GPU device assignment between tenants. Containerd 1.7.x can produce noisy-neighbor GPU issues in this configuration.
Layer 3: Network Layer (Per-Tenant CNI and Hardware-Enforced Segmentation)
The Threat Model
A default Kubernetes network is flat. Every pod can reach every other pod unless a NetworkPolicy explicitly denies the traffic. At scale, across hundreds of tenants, that model has three structural weaknesses:
- Lateral movement. A compromised pod on a flat network can scan the subnet, discover services belonging to other tenants, and attempt to exploit them. There is no Layer 2 boundary preventing it.
- Bypassable software policies. Standard Kubernetes NetworkPolicies are mutable software rules. A tenant with permission to create their own policies can inadvertently create gaps. Network policy validation is difficult at scale, and policy conflicts between tenant rules and platform rules are hard to audit.
- CNI lock-in. In a shared cluster, all tenants share the same CNI plugin. AI cloud providers serving GPU training workloads often need SR-IOV with RDMA for high-throughput interconnects. A shared CNI cannot serve both a latency-sensitive inference tenant and an RDMA-dependent training tenant on the same cluster.
The Architecture
// Weak isolation: shared CNI, software-only policies
// Traffic between tenants is permitted by default at Layer 2
Node Network (shared CNI) {
Pod-A (tenant-a, 10.0.1.5)
Pod-B (tenant-b, 10.0.1.6)
// Default: A can reach B. NetworkPolicy required to block.
// NetworkPolicy is mutable; misconfiguration opens the path.
}
// Strong isolation: per-tenant CNI + hardware VLAN segmentation
// Enforcement moves from software rules to the physical switch
Physical Switch (managed by Netris via vMetal) {
VLAN 10 (tenant-a) {
Port -> PrivateNode-1 (tenant-a worker)
Port -> PrivateNode-2 (tenant-a worker)
}
VLAN 20 (tenant-b) {
Port -> PrivateNode-3 (tenant-b worker)
}
// No Layer 2 path exists between VLAN 10 and VLAN 20.
// VRF enforces routing separation at Layer 3.
// ACLs enforce policy at the DPU if present.
}The Controls
Per-tenant CNI with Private Nodes
Because Private Nodes dedicates physical worker nodes to each tenant, each tenant cluster can run an entirely different CNI plugin. Calico with eBPF is the primary recommendation from vCluster Labs for per-tenant CNI. A second tenant on the same platform can run Multus with SR-IOV for RDMA-enabled GPU training without any conflict. One tenant can use KubeOVN for enhanced namespace-level segmentation. The per-tenant CNI model eliminates the constraint that forces all tenants onto a single shared network stack.
vMetal with Netris: hardware-enforced network segmentation
vMetal, the machine layer beneath vCluster Platform, integrates with Netris to automate network segmentation at the physical switch. It is not another provisioning tool; vMetal is a provisioning orchestrator with one stable, EC2-like API that drives Metal3, KubeVirt, NVIDIA NICo, and OpenStack under a single consistent interface.
Through the Netris integration, vMetal programmatically configures per-tenant VLANs, VXLANs, VRFs, and ACLs directly on the physical network fabric. DPU policy enforcement is also available in this path. When nodes move between tenants, the switch configuration is dynamically updated. This is the network automation layer that moves policy enforcement from mutable iptables and eBPF rules to the physical switch: a boundary that tenant workloads cannot traverse regardless of what runs at the software layer.
vCluster Platform is validated in NVIDIA's DGX reference architecture, which vCluster authored.
Composing the Full-Stack Isolation Boundary
Hardening one layer is not enough. A per-tenant control plane with a flat shared network is still vulnerable to lateral movement. Private Nodes with a shared API server still expose the control plane blast radius. Network segmentation without compute isolation leaves the kernel exposed.
The three layers compose into a single, coherent isolation boundary:
vCluster Labs ships this full-stack isolation boundary as a composed, production-ready platform. vCluster Platform provides control plane isolation. Private Nodes deliver hardware-level compute isolation as the production default. vNode adds kernel-native workload sandboxing for defense-in-depth. vMetal with Netris enforces network segmentation at the physical fabric.
This architecture is proven at scale: 100K+ GPUs powered across 50+ GPU Clouds & Fortune 500s, with 40M+ tenant clusters created. Boost Run went from decision to production in under 45 days with zero new platform engineering hires. Lintasarta launched Indonesia's leading GPU cloud in 90 days with 170+ tenant clusters in production.
Tenant isolation at this layer depth is the architecture that makes a commercial AI cloud viable for paying, untrusted external tenants.
To explore the full-stack approach, start with vCluster Platform. For the compute isolation layer specifically, see vNode.
Frequently Asked Questions
What is tenant isolation in GPU cloud infrastructure?
Tenant isolation means each tenant gets dedicated tenant clusters and separated API, compute, and network boundaries so one tenant's misconfiguration, compromise, or resource spike cannot affect another. In GPU and AI cloud infrastructure, isolation cannot stop at database controls; it must extend to the Kubernetes control plane, physical or kernel-level compute boundaries, and the network fabric. vCluster Platform composes per-tenant control planes, Private Nodes, and switch-level segmentation through vMetal and Netris into one full-stack isolation boundary.
Why are Private Nodes the default isolation model for production AI workloads?
Private Nodes give each tenant dedicated physical worker nodes and are the production default because they remove the shared-kernel blast radius that makes container escapes dangerous. On shared GPU nodes, a container escape can expose other tenants' process memory, files, and GPU state. Private Nodes ensure no other tenant's workloads share the same hardware, kernel, or OS.
How does vCluster Platform isolate the Kubernetes control plane per tenant?
vCluster Platform gives each tenant its own dedicated API server, etcd, scheduler, controllers, RBAC, admission control, and namespace-scoped CRDs inside a tenant cluster. The underlying control plane cluster remains invisible to tenants, so a misconfigured CRD or runaway controller in one tenant has zero blast radius on any other tenant cluster. This isolation also extends to Slurm, NVIDIA Run:ai, and Ray environments, with Inference and Agent Sandbox.
What is the difference between Private Nodes and vNode?
Private Nodes provide hardware-level isolation through dedicated physical worker nodes, while vNode optionally adds kernel-native sandboxing as a defense-in-depth layer. Private Nodes are the production default for enterprise and commercial GPU cloud use cases. vNode uses Linux user namespaces, cgroups, and seccomp filtering to contain a process even after a container escape. vNode is generally available and runs GPU workloads at bare-metal speed because it uses no hypervisor or guest kernel.
How does a per-tenant CNI improve network isolation?
A per-tenant CNI lets each tenant cluster run its own container network plugin, removing the forced shared network stack and enabling hardware-enforced segmentation. Because Private Nodes dedicates physical worker nodes to each tenant, one tenant can run Calico with eBPF while another runs Multus with SR-IOV for RDMA-enabled GPU training. Paired with vMetal and Netris, enforcement moves from mutable software rules to per-tenant VLANs, VXLANs, VRFs, and ACLs on the physical switch.
Why are database controls like row-level security not enough for AI tenant isolation?
Row-level security and schema separation protect data, but they cannot contain API, compute, or network compromise in GPU-accelerated AI infrastructure. AI workloads execute untrusted model-generated code, making container escape and lateral movement the primary risks. A compromised pod on a flat network or shared node can reach other tenants even when database access is properly isolated. Full isolation requires control plane, compute, and network boundaries.
What network controls does vMetal with Netris enforce?
vMetal with Netris automates per-tenant VLANs, VXLANs, VRFs, and ACLs directly on the physical switch, so tenant traffic has no Layer 2 path between isolated environments. This moves enforcement from mutable iptables and eBPF rules to the physical fabric. DPU policy enforcement is also available. When nodes move between tenants, the switch configuration updates dynamically.
Is vNode generally available for production use?
Yes, vNode is generally available and hardened for production as a kernel-native workload isolation layer. vNode uses seccomp, cgroups, and Linux namespaces to contain processes inside a sandbox, even after a container escape. It integrates with NVIDIA GPU Operator through RuntimeClass and CDI annotations and runs workloads at bare-metal speed without a hypervisor or guest kernel. For newer NVIDIA GPU Operator versions on shared nodes, containerd 2.x is required for correct GPU device assignment between tenants.
Deploy your first virtual cluster today.