Purpose-Built for GPU Node Health at Scale
Combine server health signals, Machine readiness, tenant cluster observability, and clear capacity ownership for GPU node health checks.
Observability
GPU Server Health Signals
Use BareMetalHost conditions, BMC connectivity, inspection results, and the GPU software stack to evaluate server health.
Fleet Management
Machine Readiness Across the Fleet
Track whether each Machine request has been scheduled, provisioned, joined to its target cluster, and marked ready.
Single pane of glass for all clusters
Quota and template management at scale
RBAC policies applied fleet-wide
Tenant Isolation
Private Nodes Per Tenant
Private Nodes dedicate worker capacity to one production tenant cluster at a time, making ownership of affected capacity clear.
Private Nodes with dedicated physical hardware per tenant
Per-tenant CNI and storage
Tenant-scoped capacity ownership
Dynamic Provisioning
Capacity Requests Through Auto Nodes
Auto Nodes can maintain configured Private Node capacity by creating and deleting Machine requests through vMetal.
Bare metal GPU nodes on demand
Machine requests through Auto Nodes
Reduces manual node remediation effort
Bare Metal Layer
Machine Lifecycle Visibility
vMetal exposes the lifecycle of each Machine claim and the server inventory available to satisfy future requests.
Automated Metal3 provisioning
Full machine lifecycle under one API
Rack to decommission in one workflow