Summary
- Dedicated Kubernetes clusters per customer can consume ~$12 per tenant in baseline overhead before application workloads, while hardware-isolated tenant clusters with dedicated worker nodes approach ~$5 per tenant at scale in vCluster's cost modeling.
- Hardware-level isolation (dedicated worker nodes, no shared kernel, no cross-tenant scheduling, and per-tenant CNI/CSI) can be managed from a single control plane cluster instead of provisioning a full cluster per customer.
- Use hardware-isolated tenant clusters as the production default, shared-node tenant clusters for internal dev/test, and full dedicated clusters only for true air-gap requirements.
- vCluster Platform scales this model with templates, GitOps, and self-service provisioning.
Running a dedicated Kubernetes cluster per customer means provisioning a complete, separate control plane and worker nodes for every tenant. Platform and AI cloud teams choose this model to meet hardware-level isolation requirements, satisfy compliance contracts, and eliminate noisy-neighbor contention. vCluster Private Nodes deliver the same hardware-level isolation (dedicated worker nodes per tenant, no shared kernel, no cross-tenant scheduling) without provisioning a full cluster for every customer.
Why Teams Default to One Cluster Per Customer
The reasoning is straightforward. Namespace-based isolation is logical, not physical. If an attacker compromises a node, they can reach the data of every tenant running on that node. Security and compliance reviews increasingly require proof of hardware-level isolation, not just API-level isolation.
Beyond security, there are three recurring drivers:
- Noisy-neighbor risk. A single tenant running a large model-training job can exhaust CPU, memory, or GPU capacity across shared nodes, degrading performance for every other tenant on the same hardware.
- Blast radius containment. When a shared component like Prometheus or a logging agent fails in a shared cluster, every team using that cluster is affected. In a per-customer cluster, a failure is scoped to one tenant.
- Operational familiarity. Many platform engineering teams grew up managing a one-application-per-cluster model. Extending that to one-customer-per-cluster feels like a natural fit, even if the economics do not scale.
These are valid requirements. The problem is the model used to satisfy them.
The Compounding Cost of Dedicated Clusters
The dedicated-cluster-per-customer model carries costs that do not appear in the initial architecture decision but compound with every new tenant added.
Idle resource waste. Every dedicated cluster requires a minimum viable footprint: a control plane, at least one node per availability zone, and the suite of platform add-ons (ingress, monitoring, logging, certificates). Those resources run continuously regardless of whether the customer's workload is active. At scale, cost analysis shows that 50 dedicated clusters can consume roughly $12 per tenant in baseline overhead before any application workload is counted. For illustration, a shared ingress controller might cost around $0.10 per month; multiplied across 50 dedicated instances, that ratio alone shows the scaling problem, though actual costs vary by cloud provider and configuration.
Upgrade sprawl. Each cluster must be patched, upgraded, and validated independently. At tens or hundreds of customers, this becomes unmanageable.
Day-2 operations overhead. Monitoring, log aggregation, security scanning, and access management must be configured and maintained per cluster. Tool licensing often scales per cluster, not per workload. The operational depth required to run Kubernetes reliably is already a known pain point; multiplying it by the number of customers makes it structural.
Provisioning latency. Standing up a production-ready Kubernetes cluster takes hours at minimum and often longer when accounting for DNS, certificates, networking configuration, and add-on installation. That latency translates directly into customer onboarding time, which is a constraint most growth-stage platforms cannot absorb.
Comparing Tenant Isolation Models
Three distinct approaches to tenant isolation exist in Kubernetes. They differ substantially in the guarantees they provide and the costs they impose.
Isolation ModelIsolation BoundaryProvisioning TimeCost ProfilePractical FitvCluster Private NodesHardware-level. Dedicated worker nodes per tenant. No shared kernel. No cross-tenant scheduling. Per-tenant CNI/CSI.MinutesModerate. Approaches ~$5 per tenant at scale in vCluster's cost modeling.Production default for AI platforms, SaaS providers, and compliance-required workloadsDedicated cluster per customerFull hardware and control-plane separation. True air-gap.Hours to daysHighest. Idle resource waste scales with tenant count.Rare regulatory or contractual air-gap requirementsNamespace-onlyLogical (RBAC, NetworkPolicy). Shared kernel, scheduler, CNI, and CSI. Node compromise exposes all tenants.SecondsLowestInternal dev/test environments with trusted workloads
vCluster Platform's isolation spectrum has three tiers: Shared Nodes for trusted dev/test/CI/CD, Private Nodes as the recommended production default, and Dedicated VMs + vNode for kernel-native workload isolation without hypervisor overhead. The dedicated cluster is the expensive one, not the safe default. Private Nodes occupy the position that the dedicated-cluster model was supposed to fill, at a fraction of the operational cost.
What Private Nodes Actually Are
vCluster Private Nodes give each tenant a dedicated set of worker nodes managed through a single control plane cluster. The isolation guarantee is at the hardware level. A tenant's pods run exclusively on nodes assigned to that tenant. The Kubernetes scheduler for one tenant has no visibility into and cannot schedule onto the nodes of any other tenant.
The technical properties that deliver this:
- Dedicated worker nodes. Each tenant cluster's workloads are bound to a specific node pool. Node taints and tolerations enforce this at the scheduler level. No cross-tenant pod placement is possible. Nodes are not visible from the control plane cluster; they exist only within that tenant cluster.
- No shared kernel. Because tenant workloads run on separate virtual or physical machines, there is no shared kernel between tenants. Container escape vulnerabilities that exist in namespace-based models, where a kernel exploit can reach every tenant on the node, do not apply here.
- No cross-tenant scheduling. Resource contention between tenants is eliminated at the source. A compute-intensive AI workload from one tenant cannot starve another tenant of CPU, memory, or GPU capacity.
- Per-tenant CNI and CSI. Each tenant can have its own network configuration and storage class. This allows per-customer network policies and storage isolation without requiring a separate cluster. It also simplifies the debugging of network issues, because each tenant's network stack is independently scoped.
For tenants with elevated security requirements, vCluster Platform also offers an optional VPN feature that encrypts traffic between the control plane cluster and a tenant's dedicated worker nodes. This is a separate, opt-in capability, not part of the default Private Nodes isolation guarantee.
This architecture addresses the core objection to namespace-based isolation. Node-level access does not expose other tenants because other tenants do not share the node.
When to Use Each Isolation Model
The choice depends on the workload classification and the compliance obligation, not on a general preference for isolation.
Use Private Nodes as the production default. Any workload that processes customer data, runs under a compliance framework, or carries a contractual isolation requirement belongs on Private Nodes. This covers the majority of AI cloud provider and SaaS platform use cases: model inference serving, customer data pipelines, and tenant-isolated SaaS backends. Private Nodes satisfy hardware-level isolation requirements without the cluster-per-customer overhead.
Reserve a full dedicated cluster for true air-gap requirements. Some regulated industries or government contracts require a physically and logically separate environment with its own control plane, its own network boundary, and no shared infrastructure of any kind. This is a narrow exception, not the default.
Platform teams that treat the dedicated-cluster model as the safe choice for all production workloads are paying a significant premium for a guarantee that Private Nodes provide at a fraction of the cost.
How to Offer Dedicated-per-Customer Isolation at Scale
Implementing Private Nodes at scale requires more than the node configuration itself. It requires a management layer that can provision, template, and govern tenant clusters consistently across a growing fleet.
vCluster Platform provides that layer. From a single control plane, platform teams manage all tenant clusters — whether shared-node or Private Node configurations — through a unified UI and API. This eliminates the context-switching and kubeconfig sprawl that comes with managing many independent clusters.
Templates enforce consistency. A "production-isolation" template can define Private Nodes, specific RBAC roles, network policies, and resource quotas as defaults. When a new tenant is provisioned from that template, the configuration is applied automatically. This removes the possibility of misconfiguration during onboarding and reduces the burden on the platform team.
GitOps integration automates the lifecycle. Templates integrate with ArgoCD or Flux, enabling fully automated provisioning pipelines. A new customer entry in the source repository triggers cluster creation, configuration, and deployment without manual intervention. vCluster Platform also enforces per-tenant network isolation through Netris, automating the VLAN and VXLAN boundaries that keep tenant traffic separate.
Self-service provisioning reduces coordination overhead. A self-service portal lets customers or internal teams provision their own isolated environments within pre-approved boundaries. Platform engineers define the templates; tenants consume them. This decouples onboarding speed from platform team capacity.
Behind the scenes, vMetal is the machine layer that turns raw GPU racks into the dedicated nodes Private Nodes depend on. It provisions bare metal and virtual machines through one EC2-like API, then vCluster Platform allocates those machines to tenant clusters. That gives AI cloud providers a clean path from unprovisioned hardware to customer-ready isolated environments in one stack.
The results at production scale are observable. Lintasarta manages over 170 tenant clusters through vCluster Platform. Boost Run reached production in under 45 days using vCluster for its AI cloud infrastructure.
Private Nodes are available on a free tier covering up to 32 GPUs, making it practical to evaluate hardware-level isolation at meaningful scale before committing to an enterprise plan.
Platform teams running dedicated Kubernetes clusters per customer are solving a real problem with a model that does not scale. Private Nodes provide dedicated worker nodes, no shared kernel, per-tenant CNI and CSI, and no cross-tenant scheduling (the technical substance of a dedicated cluster) managed through a single control plane at a fraction of the cost.
Request a demo to see how Private Nodes fit your current tenant isolation and compliance requirements.
Frequently Asked Questions
Is Private Nodes as secure as a dedicated cluster per customer?
For most production and compliance workloads, yes. Private Nodes provide hardware-level tenant isolation through dedicated worker nodes and no shared kernel.
Private Nodes bind each tenant cluster's workloads to a dedicated node pool. Because tenants do not share nodes, container escape and kernel-level attacks cannot cross tenant boundaries. The control plane cluster centralizes management but does not process tenant workload data directly. A fully separate dedicated cluster remains the right answer only for narrow air-gap requirements that mandate a separate control plane and network boundary.
Does Private Nodes require a separate Kubernetes cluster per customer?
No. Private Nodes use a single control plane cluster to manage all tenant clusters, while each tenant receives dedicated worker nodes.
That is the central architectural difference from a dedicated cluster per customer. One control plane cluster manages the fleet, but each tenant cluster gets its own node pool, network configuration, and storage configuration. The hardware isolation is local to each tenant; the control plane management is centralized.
How does compliance proof work at the hardware level with Private Nodes?
Auditors can verify that a tenant's workloads run only on a named set of dedicated nodes using node taints, tolerations, RBAC, and audit logs.
Node taints and tolerations enforce the placement policy at the scheduler level. RBAC and audit logs provide the evidentiary record. Demonstrating that no other tenant's pods can be scheduled onto a named node pool is straightforward: the taint configuration is explicit, and the scheduler enforces it without exception. This satisfies the most common hardware isolation requirements in enterprise and regulated-sector contracts.
What is the difference between Private Nodes and shared-node tenant clusters?
Shared-node tenant clusters isolate tenants logically on the same worker nodes and kernel. Private Nodes isolate tenants at the hardware level by giving each tenant cluster its own dedicated worker nodes.
In a shared-node model, a kernel exploit or node compromise can expose every tenant on that node. Private Nodes remove that risk by eliminating shared nodes. Shared nodes remain appropriate for internal dev/test/CI/CD environments where tenant trust is high and cost is the primary driver.
What is the cost difference between Private Nodes and dedicated clusters per customer?
Private Nodes approach roughly $5 per tenant at scale in vCluster's cost modeling, while dedicated clusters can consume around $12 per tenant in baseline overhead before application workloads are counted.
Dedicated clusters require a minimum viable footprint per customer: a control plane, at least one node per availability zone, and platform add-ons such as ingress, monitoring, logging, and certificates. Private Nodes consolidate control plane management and eliminate idle cluster-level overhead while preserving dedicated worker nodes per tenant.
Can Private Nodes support AI cloud or GPU cloud workloads?
Yes. Private Nodes are designed for AI cloud and GPU cloud use cases where noisy-neighbor contention and hardware-level tenant isolation matter.
A tenant running a large model-training job cannot starve another tenant of CPU, memory, or GPU capacity because tenants do not share nodes. Private Nodes are available on a free tier covering up to 32 GPUs, making it practical to evaluate hardware-level isolation before scaling.
How do platform teams offer Private Nodes to customers at scale?
Platform teams use templates, GitOps, and self-service provisioning through vCluster Platform to deliver tenant-isolated environments in minutes.
A production-isolation template can define Private Nodes, RBAC roles, network policies, and resource quotas as defaults. Templates integrate with ArgoCD or Flux to automate tenant cluster creation, and a self-service portal lets customers provision within approved boundaries without opening a support ticket.
When should a team still use a dedicated cluster per customer?
Use a dedicated cluster only for true air-gap requirements that mandate a separate control plane, separate network boundary, and no shared infrastructure of any kind.
Some regulated industries and government contracts require this level of physical and logical separation. For the majority of AI cloud and SaaS platforms, Private Nodes satisfy hardware-level isolation requirements without the operational and financial overhead of a dedicated cluster per customer.
Deploy your first virtual cluster today.