inference

Single-Tenant AI Inference for Sovereign Clouds

Run each inference customer in an isolated tenant cluster with its own Kubernetes control plane and Private Nodes. vCluster Platform centralizes tenant access, templates, capacity policy, and fleet operations across your GPU infrastructure.

Trusted by the fastest-growing AI cloud providers
Problem

Where Shared Inference Infrastructure Falls Short

Shared inference infrastructure needs clear tenant boundaries across the control plane, worker capacity, runtime, and network.

Shared Components Complicate Compliance

Namespace-only designs share control-plane and node-level components between customers.

Cross-Tenant Resource Contention

Shared worker nodes can expose inference workloads to cross-tenant resource contention.

Data Residency Requires Full-Stack Planning

Data residency depends on the placement and configuration of compute, storage, networking, and inference services, not on a single product setting.

Solution

Tenant Isolation for Single-Tenant AI Inference

vCluster Platform gives each inference customer a separate tenant cluster and uses Private Nodes as the production default. Optional vNode and Netris integrations strengthen the runtime and physical network boundaries.

Built for Single-Tenant AI Inference

Single-tenant AI inference combines a separate tenant control plane, Private Nodes, and optional runtime and physical network isolation.

Compute Isolation

Private Nodes for Inference Tenants

Private Nodes assign dedicated worker capacity, networking, and storage to one production tenant cluster at a time.

  • Dedicated worker capacity
  • Tenant-scoped networking and storage
  • Dedicated tenant capacity
Workload Security

Runtime Isolation With vNode

vNode uses Linux user namespaces and seccomp filters to strengthen the runtime boundary for inference workloads that need additional isolation.

  • Stronger runtime boundary
  • No additional VM layer
  • Seccomp filters
Network Isolation

Optional Hardware Network Isolation

When Metal3 and Netris are configured, separate tenant network environments can receive hardware-backed L2 isolation.

  • Per-tenant network boundaries
  • Netris network environments
  • Hardware-backed L2 isolation
Sovereign Compliance

Air-Gapped Platform Deployment

vCluster Platform supports air-gapped deployments and FIPS features on supported plans for infrastructure operated in controlled environments.

  • Air-gapped deployment supported
  • FIPS features on supported plans
  • Controlled deployment locations
Tenant Clusters

Separate Control Plane Per Tenant

Each inference tenant gets its own virtualized Kubernetes control plane, API server, and RBAC boundary.

  • Separate API server and RBAC
  • Lightweight control plane creation
  • Tenant-scoped administration

Why vCluster

This isn’t a side project. Behind every vCluster deployment is 5+ years of deep K8s engineering, security hardening, and battle-tested infrastructure work at massive scale.

100K+
GPUs Powered
50+
GPU Clouds & F500s
<45
Days to Launch
30K
GitHub Stars

Get Started in 3 Steps

1
Schedule a Demo

Talk to our team about your stack

2
Deploy vCluster

Deploy vCluster on your infra in minutes

3
Onboard Your Tenants

Go live with a hyperscaler-grade tenant experience in days

FAQs

What makes vCluster single-tenant AI inference different from namespace isolation?

Namespace-only designs share a Kubernetes API server and node-level components. vCluster gives each inference customer a separate virtualized control plane and RBAC boundary, with Private Nodes as the production default for dedicated worker capacity.

How does vCluster support data residency for sovereign AI inference?

vCluster Platform supports air-gapped deployments and FIPS features on supported plans. Data residency depends on where you deploy the platform, tenant clusters, storage, networking, and inference services, so the architecture must be configured for the required jurisdiction.

Does single-tenant inference require a separate physical cluster per tenant?

No separate Kubernetes control-plane servers are required for every customer. The tenant control plane runs as isolated pods on the control plane cluster, while Private Nodes provide dedicated production worker capacity.

How does vNode strengthen inference workload isolation?

vNode is a tenant isolation container runtime that uses Linux user namespaces and seccomp filters to strengthen the workload boundary. It complements the tenant control plane and Private Nodes without taking responsibility for inference scheduling or GPU allocation.

Can vCluster support inference tenants across regions or data centers?

Yes. One vCluster Platform instance can manage connected control plane clusters across regions and infrastructure environments. Operators retain central control of access, templates, capacity, and observability.

How quickly can providers create isolated inference environments?

Virtualized control planes are lightweight and can be created quickly. Provisioning time for Private Nodes depends on whether capacity is already available or must be provisioned through a cloud or bare metal driver.

Launch Isolated Inference Environments on Your GPU Hardware

See how vCluster Platform delivers single-tenant AI inference environments with hard tenant isolation.