Built for Resilient GPU Infrastructure at Scale
Kubernetes node autohealing depends on cluster health signals, tenant-scoped capacity, Machine lifecycle automation, and infrastructure repair tooling.
Automated Operations
Centralized Node Health Operations
Use vCluster Platform observability alongside infrastructure health signals to coordinate node recovery across tenant clusters.
Cluster and Machine health context
Tenant cluster observability
Infrastructure health signals
Hardware Isolation
Private Nodes Per Tenant Cluster
Private Nodes assign dedicated worker capacity to one production tenant cluster at a time.
Per-tenant Private Nodes with workload isolation
Per-tenant CNI and storage
Clear tenant capacity ownership
Dynamic Provisioning
Machine Requests for Replacement Capacity
Auto Nodes can create and delete Machine requests to maintain configured Private Node capacity through vMetal.
Fleet Operations
Fleet-Wide Tenant Cluster Operations
Manage tenant clusters through one UI, CLI, and API with templates, access, policy, and fleet observability.
Tenant Isolation
Isolated Control Planes Per Tenant
Each tenant cluster has its own virtualized control plane and RBAC boundary, separate from worker node recovery operations.
Own API server per tenant
Failure scoped to one tenant
Recovery separate from control plane