ai-cloud

GPU Server Provisioning for AI Clouds

vMetal exposes bare metal servers and VMs through one stable Machine API. Interchangeable compute and network drivers handle provisioning and lifecycle operations, while vCluster turns that capacity into tenant cluster products.

Trusted by the fastest-growing AI cloud providers
Problem

Why GPU Provisioning Breaks at Scale

GPU server provisioning becomes difficult when every infrastructure stack exposes a different lifecycle and API.

Race to the Bottom

Selling raw compute alone gives customers no consistent machine or tenant cluster experience.

Slow, Manual Provisioning

Manual OS installation, network setup, and cluster registration slow the path from available hardware to usable capacity.

Compliance Without a Path

Air-gapped and regulated environments need machine provisioning that runs on infrastructure the operator controls.

Solution

One API From Raw Rack to Tenant Cluster

vMetal presents one stable Machine API for physical and virtual servers. Infrastructure drivers handle provisioning, and released capacity returns to inventory for reuse.

Full Stack GPU Server Provisioning

vMetal provides the Machine API and lifecycle layer. vCluster consumes those Machines as tenant cluster capacity.

Machine Layer

Automated GPU Server Provisioning

With the Metal3 driver, vMetal provisions a bare metal server, installs the selected OS image, applies network configuration, and tracks the Machine lifecycle through one stable API.

  • PXE boot to production automatically
  • One API across all provisioning stacks
  • Full GPU server lifecycle management
Dynamic Scaling

Machine Requests for Private Nodes

Auto Nodes can create and delete Machine requests to maintain Private Node capacity for tenant clusters.

  • GPU nodes provisioned on workload schedule
  • Capacity requests through Auto Nodes
  • Release capacity when tenants are idle
Tenant Isolation

Private Nodes for Tenant Capacity

vCluster consumes provisioned Machines as Private Nodes, assigning dedicated worker capacity to one production tenant cluster at a time.

  • Private Nodes per tenant
  • Per-tenant CNI and storage stack
  • No cross-tenant workload exposure
Network Security

Optional Hardware Network Isolation

When Netris is configured with Metal3, separate network environments can receive hardware-backed L2 isolation.

  • VLANs and VRFs per tenant
  • Network automation through supported integrations
  • Hardware-backed L2 isolation with Netris
Compliance

Air-Gapped Platform Deployment

vCluster Platform supports air-gapped deployments and FIPS features on supported plans for infrastructure operated in controlled environments.

  • Air-gapped deployment supported
  • FIPS-capable configuration available
  • Meets data residency mandates

Why vCluster

This isn’t a side project. Behind every vCluster deployment is 5+ years of deep K8s engineering, security hardening, and battle-tested infrastructure work at massive scale.

100K+
GPUs Powered
50+
GPU Clouds & F500s
<45
Days to Launch
30K
GitHub Stars

Get Started in 3 Steps

1
Schedule a Demo

Talk to our team about your stack

2
Deploy vCluster

Deploy vCluster on your infra in minutes

3
Onboard Your Tenants

Go live with a hyperscaler-grade tenant experience in days

FAQs

What is vMetal and how does it simplify GPU server provisioning?

vMetal is the physical and virtual Machine provisioning and lifecycle layer of vCluster Platform. A Machine is a claim on one server, and the selected driver handles the infrastructure-specific provisioning work behind a stable API.

Can vMetal handle provisioning for sovereign or air-gapped GPU clouds?

vMetal is designed for infrastructure that the operator controls, including on-premises bare metal. Air-gapped and FIPS requirements should be implemented through the supported vCluster Platform deployment options and the selected infrastructure drivers.

What sellable products does GPU server provisioning with vMetal produce?

vMetal provisions Bare Metal Machines and Virtual Machines. vCluster consumes that capacity through Private Nodes and turns it into managed Kubernetes, Slurm, Ray, inference, and other tenant cluster products.

How does Auto Nodes work for dynamic GPU server provisioning?

Auto Nodes is a vCluster consumption workflow. It creates and deletes Machine requests through vMetal to maintain configured Private Node capacity based on static counts or supported resource requirements.

How long does it take to go from raw GPU racks to tenant-ready infrastructure?

Provisioning time depends on server state, the selected driver, OS image, network setup, and startup automation. vMetal keeps the request and lifecycle model consistent across those infrastructure choices.

What network isolation options are available for provisioned GPU servers?

With the Metal3 and Netris integration configured, separate network environments can provide hardware-backed L2 isolation. Metal3 without that integration can also run on a flat network, so the network mode must be selected deliberately.

Turn Your GPU Racks Into a Cloud

See how vMetal and vCluster handle GPU server provisioning end to end.