ai-cloud

GPU Node Lifecycle Management for AI Clouds

Manage GPU node supply through vMetal's stable Machine API. Track each request from scheduling and provisioning through cluster attachment, deletion, cleaning, and return to available inventory.

Trusted by the fastest-growing AI cloud providers
Problem

Why GPU Node Ops Break at Scale

GPU node operations become fragmented when provisioning, cluster attachment, and inventory use different lifecycle models.

Manual Provisioning Kills Margins

Manual provisioning delays the point when GPU servers can accept workloads.

No Unified API Across Hardware

Different infrastructure drivers expose different mechanics and operational states.

Decommissioning and Updates Cause Downtime

Without a clear reclaim workflow, released servers can remain stranded outside the available capacity pool.

Solution

One API for the Full GPU Node Lifecycle

vMetal provides one Machine model across physical and virtual servers. vCluster requests capacity through Private Nodes, and released servers return to the pool for reuse.

Full GPU Node Lifecycle in One Platform

vMetal provides a consistent Machine lifecycle across requests, drivers, tenant cluster attachment, release, and inventory reuse.

Bare Metal Ops

GPU Machine Provisioning

Infrastructure drivers provision physical or virtual servers behind the stable vMetal Machine API.

  • PXE boot and OS install automated
  • Machine registration and network config
  • From rack to sellable inventory
Dynamic Scaling

Private Node Capacity Requests

Auto Nodes can create and delete Machine requests to maintain the configured Private Node capacity for tenant clusters.

  • Automatic GPU node provisioning
  • Demand-driven bare metal scaling
  • No manual node lifecycle steps
Tenant Isolation

Dedicated Tenant Capacity

Provisioned Machines can join one tenant cluster at a time as Private Nodes with dedicated worker capacity.

  • Private GPU nodes per tenant
  • Per-tenant CNI and storage
  • Workload isolation at the infrastructure boundary
Network Isolation

Network Configuration by Driver

The selected driver applies network configuration. Metal3 can integrate with Netris for hardware-backed L2 isolation when configured.

  • Automated VLAN and VXLAN setup
  • Network automation per node
  • Hard network isolation per tenant
Lifecycle Operations

Inventory and Capacity Reclaim

When a Machine is deleted, the driver cleans the server and returns it to available inventory for another claim.

  • Fleet-wide GPU node observability
  • Automated updates and backups
  • Disaster recovery and compliance

Why vCluster

This isn’t a side project. Behind every vCluster deployment is 5+ years of deep K8s engineering, security hardening, and battle-tested infrastructure work at massive scale.

100K+
GPUs Powered
50+
GPU Clouds & F500s
<45
Days to Launch
30K
GitHub Stars

Get Started in 3 Steps

1
Schedule a Demo

Talk to our team about your stack

2
Deploy vCluster

Deploy vCluster on your infra in minutes

3
Onboard Your Tenants

Go live with a hyperscaler-grade tenant experience in days

FAQs

What does GPU node lifecycle management cover in vMetal?

vMetal tracks the Machine lifecycle from request and scheduling through provisioning, cluster attachment, deletion, and cleanup. The server remains in inventory and can be claimed again after the driver returns it to the available pool.

How does vMetal handle node lifecycle management across different hardware vendors?

The Machine API stays consistent while infrastructure-specific drivers handle physical or virtual server provisioning. Each driver publishes its own requirements, capabilities, and Day 2 operations.

Can vMetal automatically scale GPU nodes based on tenant workload demand?

Auto Nodes is a vCluster workflow that maintains Private Node capacity by creating and deleting Machine requests through vMetal. Dynamic behavior follows supported resource requirements such as reserved or used CPU and memory.

How does GPU node lifecycle management integrate with tenant cluster orchestration?

vMetal owns Machine provisioning and lifecycle. vCluster Platform owns machine types, access, and quotas, while vCluster requests and releases capacity for tenant clusters through Private Nodes and Auto Nodes.

What network automation is included in GPU node lifecycle management?

Network behavior depends on the selected driver. With Metal3 and Netris configured, separate network environments can provide hardware-backed L2 isolation.

Is vMetal GPU node lifecycle management suitable for sovereign AI cloud environments?

vMetal runs on infrastructure the operator controls. Air-gapped and FIPS requirements should use the supported vCluster Platform deployment options and documented infrastructure driver requirements.

Automate Your GPU Node Lifecycle Today

See how vMetal turns raw GPU racks into sellable, managed inventory.