ai-cloud

Reserved GPU Capacity for AI Clouds

Offer reserved GPU capacity through Private Nodes assigned to one tenant cluster at a time. vCluster Platform manages tenant access and capacity policy, while vMetal can supply bare metal Machines when physical capacity must be provisioned.

Trusted by the fastest-growing AI cloud providers
Problem

Why Reserved GPU Capacity Is Hard

Reserved GPU services need clear ownership of worker capacity and a consistent way to deliver it to tenant clusters.

Separate Physical Clusters Add Operational Work

A separate Kubernetes management stack for every tenant increases infrastructure and operational work.

Namespace Isolation Does Not Reserve Nodes

Namespace-only designs do not dedicate worker capacity or provide a separate tenant control plane.

Custom Capacity Workflows Add Handoffs

Custom capacity workflows create manual handoffs between tenant requests, Machine provisioning, and cluster attachment.

Solution

One Platform to Deliver Reserved GPU Capacity

vCluster Platform assigns Private Nodes to production tenant clusters and applies tenant access and capacity policy centrally. vMetal can provision physical Machines when additional bare metal capacity is required.

Built for Reserved GPU Capacity

Reserved GPU capacity combines dedicated Private Nodes, tenant control planes, capacity policy, and optional Machine provisioning.

Tenant Isolation

Private Nodes for Reserved GPU Capacity

Private Nodes dedicate worker capacity, networking, and storage to one production tenant cluster at a time.

  • Dedicated physical nodes per tenant
  • Per-tenant CNI and storage
  • Dedicated tenant capacity
Control Plane

Separate Tenant Control Planes

Each tenant receives its own virtualized Kubernetes control plane, API server, and RBAC boundary.

  • Own API server and etcd per tenant
  • Lightweight control plane creation
  • No separate control-plane servers
Dynamic Provisioning

Capacity Requests Through Auto Nodes

Auto Nodes can create and delete Machine requests to maintain configured Private Node capacity based on supported resource requirements.

  • Machine requests through Auto Nodes
  • Configured capacity policy
  • vMetal Machine lifecycle
Bare Metal Layer

Bare Metal Machines Through vMetal

vMetal provisions physical Machines behind a stable API and can join completed servers to the target tenant cluster automatically.

  • Automated Metal3 provisioning
  • Stable Machine API
  • Full Machine lifecycle management
Platform Operations

Central Tenant and Capacity Management

Manage tenants, access, templates, quotas, capacity policy, and fleet observability through one UI, CLI, and API.

  • Central UI, CLI, and API
  • Quota and capacity management
  • SSO and cluster templates included

Why vCluster

This isn’t a side project. Behind every vCluster deployment is 5+ years of deep K8s engineering, security hardening, and battle-tested infrastructure work at massive scale.

100K+
GPUs Powered
50+
GPU Clouds & F500s
<45
Days to Launch
30K
GitHub Stars

Get Started in 3 Steps

1
Schedule a Demo

Talk to our team about your stack

2
Deploy vCluster

Deploy vCluster on your infra in minutes

3
Onboard Your Tenants

Go live with a hyperscaler-grade tenant experience in days

FAQs

How does vCluster Platform provide reserved GPU capacity per tenant?

Private Nodes assign dedicated worker capacity to one production tenant cluster at a time. Each tenant also gets its own virtualized control plane and RBAC boundary, while capacity policy controls which node types the tenant can request.

Do I need separate physical clusters for each tenant's reserved GPU capacity?

No separate Kubernetes control-plane servers are required for each tenant. Private Nodes provide dedicated worker capacity, and the tenant control plane runs as isolated pods on the control plane cluster.

What cluster types can use reserved GPU capacity?

Reserved capacity can back managed Kubernetes, Slurm, Run:AI, Ray, inference, and other cluster products. vCluster provides the tenant boundary, while Slurm and Run:AI remain responsible for their scheduling capabilities.

How quickly can I launch reserved GPU capacity for clients?

Launch time depends on whether the reserved nodes already exist or must be provisioned. vCluster provides the tenant cluster and capacity workflow, while the selected Machine driver determines infrastructure provisioning time.

How does vCluster handle network isolation for reserved GPU capacity?

Private Nodes provide tenant-scoped networking. When Metal3 and Netris are configured, separate network environments can also provide hardware-backed L2 isolation.

What evidence supports vCluster for GPU capacity at scale?

vCluster powers 100K GPUs across 50+ GPU Clouds & Fortune 500s and is validated in the NVIDIA DGX reference architecture.

Start Offering Reserved GPU Capacity Today

See how vCluster Platform supports reserved gpu capacity for ai clouds.