GPU Node Lifecycle Management for AI Clouds
Manage GPU node supply through vMetal's stable Machine API. Track each request from scheduling and provisioning through cluster attachment, deletion, cleaning, and return to available inventory.
Manage GPU node supply through vMetal's stable Machine API. Track each request from scheduling and provisioning through cluster attachment, deletion, cleaning, and return to available inventory.
GPU node operations become fragmented when provisioning, cluster attachment, and inventory use different lifecycle models.
Manual provisioning delays the point when GPU servers can accept workloads.
Different infrastructure drivers expose different mechanics and operational states.
Without a clear reclaim workflow, released servers can remain stranded outside the available capacity pool.
vMetal provides one Machine model across physical and virtual servers. vCluster requests capacity through Private Nodes, and released servers return to the pool for reuse.
vMetal provides a consistent Machine lifecycle across requests, drivers, tenant cluster attachment, release, and inventory reuse.
Infrastructure drivers provision physical or virtual servers behind the stable vMetal Machine API.

Auto Nodes can create and delete Machine requests to maintain the configured Private Node capacity for tenant clusters.

Provisioned Machines can join one tenant cluster at a time as Private Nodes with dedicated worker capacity.

The selected driver applies network configuration. Metal3 can integrate with Netris for hardware-backed L2 isolation when configured.

When a Machine is deleted, the driver cleans the server and returns it to available inventory for another claim.

This isn’t a side project. Behind every vCluster deployment is 5+ years of deep K8s engineering, security hardening, and battle-tested infrastructure work at massive scale.
Talk to our team about your stack
Deploy vCluster on your infra in minutes
Go live with a hyperscaler-grade tenant experience in days
vMetal tracks the Machine lifecycle from request and scheduling through provisioning, cluster attachment, deletion, and cleanup. The server remains in inventory and can be claimed again after the driver returns it to the available pool.
The Machine API stays consistent while infrastructure-specific drivers handle physical or virtual server provisioning. Each driver publishes its own requirements, capabilities, and Day 2 operations.
Auto Nodes is a vCluster workflow that maintains Private Node capacity by creating and deleting Machine requests through vMetal. Dynamic behavior follows supported resource requirements such as reserved or used CPU and memory.
vMetal owns Machine provisioning and lifecycle. vCluster Platform owns machine types, access, and quotas, while vCluster requests and releases capacity for tenant clusters through Private Nodes and Auto Nodes.
Network behavior depends on the selected driver. With Metal3 and Netris configured, separate network environments can provide hardware-backed L2 isolation.
vMetal runs on infrastructure the operator controls. Air-gapped and FIPS requirements should use the supported vCluster Platform deployment options and documented infrastructure driver requirements.
See how vMetal turns raw GPU racks into sellable, managed inventory.