Tech Blog by vClusterPress and Media Resources

vCluster Auto Nodes: Provision Azure Workers From an EKS-Hosted Control Plane

vCluster Auto Nodes: Provision Azure Workers From an EKS-Hosted Control Plane
Auto Nodes multi-cloud concept

What is Auto Nodes?

  • vCluster Auto Nodes gives each tenant cluster real, dedicated nodes created on demand, not carved out of a shared pool.
  • When a pod goes pending, Terraform provisions a real cloud instance (EC2, Azure VM, or GCE) and attaches it to the tenant cluster.
  • When the node is no longer needed, it is released back to the cloud provider — you don't pay for idle compute capacity.
  • vCluster does all the heavy lifting, so the same pattern works identically on AWS, Azure, GCP, and other providers.
Auto Nodes lifecycle

Auto Nodes provides one clear guarantee: every node a tenant cluster receives is dedicated to that tenant. Nodes are claimed only when needed and released when they are no longer required. They are never shared with another tenant, and the same guarantee applies across AWS, Azure, and GCP.

This Article Walks Through

  • A Control Plane Cluster running on AWS EKS.
  • A Node Provider configured for Azure, claiming real Azure VMs on demand.
  • This combination is deliberate: it shows that the cloud hosting vCluster Platform and the cloud supplying tenant nodes don't have to match. You can mix and claim nodes from any supported provider, regardless of where the Control Plane Cluster runs.
  • vCluster supports Auto Nodes on AWS, Azure, and GCP with the same steps. This article walks through Azure end to end — see References for the AWS and GCP repos.

Prerequisites

Local tools

  • kubectl and access to a Kubernetes cluster
  • The vCluster Platform CLI / Helm access to install Platform

A host cluster

  • An existing Kubernetes cluster to run vCluster Platform on. This walkthrough uses AWS EKS, but any Kubernetes cluster works.

Cloud accounts

  • AWS account with an EKS cluster (used to host vCluster Platform).
  • Azure subscription (used to claim tenant nodes). This is the only cloud account you strictly need for the node-claiming part of this walkthrough.
  • GCP works the same way if you want to add it later see the vcluster-auto-nodes-gcp repo.

Architecture

Auto Nodes architecture: an AWS EKS-hosted Control Plane Cluster running vCluster Platform with the ms-azure Node Provider, provisioning VMs in an Azure cloud account that join the tenant cluster as nodes

Note: the host cluster's cloud and the node's cloud don't need to match. aws-ec2 and gcp-compute Node Providers work exactly the same way. A single tenant cluster can even mix providers.

Step 1: Access the Platform

  • Get access to the EKS host cluster (kubectl configured against it).
  • Install vCluster Platform onto the EKS cluster.
  • Log in to the Platform UI.

From here, we'll set up Auto Nodes on Azure. Platform itself runs on AWS EKS, while the tenant cluster's nodes come from Azure — the host cluster's cloud and the node's cloud don't have to match.

Step 2: Auto Nodes on Azure

2.1 Resource Group and Service Principal

Azure's Node Provider needs an existing resource group. It does not create one for you. It also needs a Service Principal scoped to that resource group:

Empty Azure resource group before Auto Nodes creates resources in it

Attach the permissions from the repo's auto_nodes_role.json to the Service Principal as a custom role scoped to that resource group. This allows Auto Nodes to create the required resources in your Azure subscription. Then generate credentials for the Service Principal, including appId, password, and tenant. You will provide them when creating the Node Provider.

2.2 Create the Azure Node Provider

In the Platform UI, go to Infra Providers → Create Infra Provider and use the Quickstart flow for Azure. Under Authentication Method, select Specify credentials inline, then provide the appId, password, and tenant generated in the previous step as the ARM_CLIENT_ID, ARM_CLIENT_SECRET, and ARM_TENANT_ID environment variables, plus your ARM_SUBSCRIPTION_ID. These credentials allow the Node Provider to create nodes in your Azure subscription.

Azure Node Provider Quickstart panel with Specify credentials inline selected and the ARM credential environment variables

After it is created, both clouds appear side by side, each with its own node types and cost:

AWS and Azure Node Providers listed side by side in the Platform UI

2.3 Create the Tenant Cluster

In the Platform UI, go to Tenant Clusters → Create Tenant Cluster and choose the Azure Node Provider. Add the autoNodes block to the YAML editor:

Tenant cluster YAML editor with the autoNodes block for Azure

privateNodes:
 enabled: true
 autoNodes:
   - provider: ms-azure
     dynamic:
       - name: az-cpu-nodes
         location: eastus
         resource-group: rg-autonode-test-sWUxeW
         nodeTypeSelector:
           - property: instance-type
             operator: In
             values: ["Standard_D2s_v5", "Standard_D4s_v5", "Standard_D8s_v5"]
         limits:
           cpu: "100"
           memory: "200Gi"

Click Create.

2.4 Success

Running Tenant Kubernetes v1.36.0
az-private-auto-node-8eeae3a5   Ready   Worker

Azure node showing Ready status in the Platform UI

In the Azure Portal, the resource group contains exactly what this pattern should produce: the VM, its NIC and disk, a NAT Gateway with its own Public IP, a managed identity, a VNet, and one network security group.

Azure Portal resource group showing the resources Auto Nodes created

Note: the "specify credentials inline" step can sometimes fail silently, leaving the Node Provider without valid credentials. If nodes don't claim, double check the Service Principal secret was saved correctly.

Note: some subscriptions have a VM family quota of 0 for the instance types above. If node claims fail, request a quota increase for that VM family in the target region first.

2.5 Confirm It: A Workload That Actually Claims a Node

A tenant cluster with one Ready node doesn't fully prove Auto Nodes reacts to real demand — that first node may just be there to run the tenant cluster's own system pods. To see a workload directly trigger a new claim, deploy something and then scale past what the existing node can hold.

apiVersion: apps/v1
kind: Deployment
metadata:
 name: test-workload
spec:
 replicas: 1
 selector:
   matchLabels:
     app: test-workload
 template:
   metadata:
     labels:
       app: test-workload
   spec:
     containers:
     - name: nginx
       image: nginx:1.27
       resources:
         requests:
           cpu: "500m"
           memory: "512Mi"
         limits:
           cpu: "1"
           memory: "1Gi"

kubectl apply -f test-workload.yaml
kubectl scale deployment test-workload --replicas=5

Five replicas at 500m CPU each is more than the existing node has free. Two pods stay Pending, and a second entry shows up under Nodes → Node Claims within seconds:

NAME         STATUS    NODE CLAIM TYPE   NODE TYPE                   DESIRED CAPACITY
auto-vxc6s   Pending   Dynamic Pool      ms-azure.standard-d2s-v3    1.18 CPU

Node Claims tab showing a new claim in Pending status alongside the existing Available node

Opening the claim shows Terraform actively provisioning the new node:

Node claim detail showing Scheduled true, Provisioned false, Joined false, with Terraform running in the provisioning logs

A minute or two later, Terraform finishes and a second Azure VM joins as a node:

NAME            STATUS   ROLES
auto-bb2555e3   Ready    <none>
auto-ae7797ce   Ready    <none>

Node claim detail showing Scheduled true, Provisioned true, Joined true, with Terraform completed successfully

And all five replicas end up Running:

test-workload-6fccf564d4-zk7ch   1/1   Running
test-workload-6fccf564d4-wwp2k   1/1   Running
test-workload-6fccf564d4-fh2nh   1/1   Running
test-workload-6fccf564d4-9266n   1/1   Running
test-workload-6fccf564d4-5tlv8   1/1   Running

Pods list showing all five test-workload replicas Running with 1/1 Ready

That's the full loop, end to end: a pod goes Pending → Auto Nodes opens a claim → Terraform provisions a real Azure VM → the VM joins as a node → the pod schedules.

Networking, at a glance

  • AWS: a NAT Gateway per subnet, giving each private node outbound internet access without a public IP.
  • Azure: a NAT Gateway with its own Public IP, attached to the subnet Auto Nodes creates.
  • GCP: a Cloud Router paired with Cloud NAT (source_subnetwork_ip_ranges_to_nat = ALL_SUBNETWORKS_ALL_IP_RANGES), so compute instances get outbound access with no public IP.

Readiness Check

Once the tenant cluster reaches Running, the Azure node shows up as Ready in both the Platform UI and kubectl get nodes. The node is a real Azure VM, dedicated to this tenant cluster, and visible in the Azure Portal alongside the other resources Auto Nodes created for it.

Why This Pattern Matters

  • Less idle compute: Dynamic nodes exist only while the tenant cluster needs them.
  • No cross-tenant sharing every node is dedicated to the tenant cluster that claimed it.
  • No manual cloud console work Terraform handles VM, networking, and identity setup behind the Node Provider.
  • Works the same way across providers, so teams aren't locked into a single cloud's node lifecycle.
  • We walked through Azure here, but the same steps claim nodes on AWS and GCP too see the vcluster-auto-nodes-aws and vcluster-auto-nodes-gcp repos.

References

Share:
Get started with the
#1 platform for AI infra.

Trusted by today’s fastest-growing AI cloud builders.

Ready to take vCluster for a spin?

Deploy your first virtual cluster today.