Tech Blog by vClusterPress and Media Resources

Standing Up NVIDIA Run:ai as a Repeatable Stack: Introducing vCluster Stacks

Oct 6, 2026
|
12
min Read
Standing Up NVIDIA Run:ai as a Repeatable Stack: Introducing vCluster Stacks

Getting a Kubernetes cluster running is only the beginning of an AI platform. Before a team can submit its first GPU workload, somebody still has to connect the registry credentials, ingress, certificates, GPU components, control plane, and cluster registration. Each piece has its own configuration, and several depend on values that do not exist until an earlier step finishes.

That is where installation guides start turning into platform engineering work. Which component needs to be healthy first? Where does its output go? What should we inspect when the next installation is waiting? And how do we repeat the same setup for another team without rediscovering those answers?

vCluster Platform 4.12 includes Stacks for describing that workflow. In this walkthrough, we will use the NVIDIA Run:ai Stack to create a tenant cluster, follow its deployment, and open NVIDIA Run:ai. We will then look at how to adapt the Stack and use the resulting platform to create projects and run workloads.

What a Stack actually manages

A StackTemplate describes a group of applications, their dependencies, configuration parameters, and outputs. A StackInstance applies that template to a particular tenant cluster or control plane cluster. Individual tasks create AppInstances or ArgoCDApplications.

The useful part is the relationship between the tasks. A task with dependsOn waits until its dependencies report healthy. Tasks without dependencies can proceed concurrently. A task can also capture an output, such as an ingress address or a value from a Secret, for a later task to consume.

For NVIDIA Run:ai, this means the deployment can obtain the ingress address before configuring the endpoint, wait for the control plane before registering the cluster, and pass the registration credentials into the cluster components. These relationships live in a reusable definition. See the Stack template documentation for the resource model and task syntax.

Selected task dependencies in the dedicated NVIDIA Run:ai Stack, where registration follows the control plane and cluster components wait for registration credentials and the GPU task

Figure 1: Selected dependencies in the dedicated NVIDIA Run:ai Stack. Registration follows the control plane, and cluster components wait for both registration credentials and the GPU task. The live template contains the complete graph.

Why NVIDIA Run:ai is a useful example

Deploying NVIDIA Run:ai involves several kinds of work: creating Kubernetes resources, installing charts, waiting for infrastructure, and calling the NVIDIA Run:ai API. A successful Helm command only covers part of that sequence.

Without a shared workflow, an engineer has to keep track of those dependencies across commands and values files. With a Stack, the platform team defines the sequence once and supplies environment-specific inputs for each deployment. When something waits or fails, the task status gives us a place to start investigating.

This is most useful when several teams need the same platform components, or when the setup includes generated values and readiness dependencies. A single independent chart may not need that extra structure. Stacks also do not supply missing GPU capacity, registry access, storage, or a production certificate. We still need to make those choices.

The NVIDIA Run:ai integration supports dedicated and central control-plane models. We will follow the dedicated model used in the screenshots: each tenant gets its own NVIDIA Run:ai control-plane lifecycle. The central model shares that control plane and manages tenant registration separately.

Architecture of the demo: a tenant Kubernetes API with NVIDIA Run:ai inside it and NVIDIA Run:ai projects organizing workloads on shared physical GPU workers

Figure 2: The demo uses a tenant Kubernetes API with NVIDIA Run:ai inside it. NVIDIA Run:ai projects organize workloads within that tenant. Physical GPU workers are shared in this configuration; creating a dedicated NVIDIA Run:ai control plane does not create private worker nodes.

To note: Shared nodes suit trusted internal tenants. Workloads still share the underlying nodes and kernel. Use a deployment designed around private nodes when the required tenant boundary includes dedicated workers. Review the shared-node security guidance before choosing this model.

Before you begin

Have the following available:

  • A Kubernetes cluster with GPU-capable nodes, sufficient CPU and memory for the platform services, and administrator access. Here we use GKE. On a default GKE GPU pool, Google manages the driver and device plugin, so we do not need to take care of them.
  • The vCluster CLI, kubectl, and vCluster Platform 4.12. Use a compatible vCluster 0.37 release.
  • A working StorageClass, LoadBalancer support, and DNS resolution for the NVIDIA Run:ai endpoint.
  • NVIDIA Run:ai registry credentials that can pull its images, plus an administrator email and password for the NVIDIA Run:ai UI.
  • A clear owner for the node-level NVIDIA driver, runtime configuration, and device plugin.

Step 1: Start vCluster Platform

Begin in the context of the Kubernetes cluster that will host Platform. Confirm that context and inspect the available storage before installing anything:

kubectl get storageclass

Then install Platform:

vcluster platform start

The captured installation completed and printed a Platform URL, along with the initial administrator credentials. Open the URL returned by your own installation and sign in. Those credentials belong to vCluster Platform; the NVIDIA Run:ai administrator account is configured later:

vCluster Platform was successfully installed and can now be reached at: https://4h8a3nz.loft.host

Thanks for using vCluster Platform!

23:28:49 done You are successfully logged into vCluster Platform!
- Use `vcluster platform create vcluster` to create a new virtual cluster
- Use `vcluster platform add vcluster` to add an existing virtual cluster to a vCluster platform instance

The successful startup confirms that Platform is reachable. It does not yet tell us anything about NVIDIA Run:ai or GPU scheduling. If Platform is already installed, use that installation and move on to the bundled templates.

Note the StorageClass name from the first command. You will need it in Step 3.

Step 2: Inspect the NVIDIA Run:ai Stack

Open Management > Stacks & Apps and select Stacks. The catalog includes the dedicated NVIDIA Run:ai Stack and the components used by the central model.

The Stacks catalog in vCluster Platform listing the NVIDIA Run:ai deployment models

Figure 3: The Stacks catalog exposes the available NVIDIA Run:ai deployment models. The demo continues with the dedicated control-plane Stack.

Open run-ai-dedicated-control-plane and inspect its graph. The captured view shows tasks for namespaces, registry credentials, ingress, bootstrap, Prometheus, GPU components, and the NVIDIA Run:ai backend. Registration and cluster installation continue further along the graph.

Task graph of the dedicated NVIDIA Run:ai Stack showing which tasks depend on earlier work

Figure 4: The graph shows which tasks depend on earlier work. For example, the backend needs registry and bootstrap configuration, while authentication waits for the backend.

It is worth looking at this before creating the tenant. If a task later stops progressing, we can trace what it is waiting for instead of treating the whole installation as one failed command. The bundled NVIDIA Run:ai Stacks use App tasks, so this demo does not require an Argo CD integration.

Step 3: Create the tenant from its template

Next, open Management > Templates > Tenant Clusters. Find NVIDIA Run:ai tenant, whose resource name is runai-tenant. This VirtualClusterTemplate creates the tenant and attaches the dedicated Stack through deploy.stacks.

The NVIDIA Run:ai tenant template in the Tenant Clusters templates list

Figure 5: NVIDIA Run:ai tenant is the tenant-cluster entry point. The StackTemplate describes the applications; this template creates the cluster that will receive them.

In the intended Platform project, create a tenant cluster and select the NVIDIA Run:ai tenant template. Use runai as the tenant-cluster ID for this example, choose an owner with access to the destination, and select the control plane host.

The tenant cluster creation form with the NVIDIA Run:ai tenant template and host selected

Figure 6: The creation form connects a tenant cluster to the selected template and host. Keep the distinction between the tenant-cluster ID and the NVIDIA Run:ai tenant-name parameter visible.

Fill in the template parameters using values from your environment:

  • Tenant name: the NVIDIA Run:ai tenant identifier. The captured form uses runai-tenant, while the vCluster ID is runai.
  • Storage class: a class that exists on the host. The screenshot uses standard-rwo; do not assume that name exists on another provider, and use the one you found in Step 1.
  • Image registry: use the server, username, token, and email supplied for your NVIDIA Run:ai registry access. These authenticate image pulls, not UI sign-in. See the NVIDIA Run:ai preparation instructions.
  • Control-plane administrator: choose the email and password for the initial NVIDIA Run:ai account.
  • Ingress provider: use the setting that matches the LoadBalancer. The standard path expects an IPv4 address; the AWS path handles the corresponding load-balancer hostname.
  • Control-plane domain: use your own domain when required. In the IPv4 demo path, leaving it empty allows a hostname to be derived with nip.io after ingress receives an address.
  • Tenant image pull Secret: provide one if the Platform image requires it in your environment.
The template parameters form collecting storage class, registry access, administrator credentials, and ingress settings

Figure 7: The demo form collects storage, registry access, administrator credentials, and ingress settings. Use your own values; the masked fields are required deployment inputs.

Step 4: Check the Stack and the tenant separately

Platform creates the StackInstance for the tenant. Start with its task status through the Platform management API:

vcluster platform connect management
kubectl get stackinstances -A

The tenant status page provides another view:

The tenant status page showing the tenant as Running with a RunAI_UI link

Figure 8: The captured tenant is Running and has a RunAI_UI link. The graph reports synced pods and a running count, but this page alone does not establish that every application is ready or that a GPU workload has executed.

Step 5: Open NVIDIA Run:ai

Use the RunAI_UI link on the tenant page, if present, or the endpoint published by your deployment. Sign in with the NVIDIA Run:ai administrator account configured in Step 3.

The NVIDIA Run:ai sign-in screen reached through the deployed ingress

Figure 9: The NVIDIA Run:ai sign-in screen is reachable through the deployed ingress. It uses the NVIDIA Run:ai account, separate from the Platform login.

Note: The captured demo uses a self-signed certificate, which explains the browser's certificate warning. For a shared or production endpoint, configure a domain you control and an appropriate trusted certificate. Do not carry a demo certificate bypass into the production instructions.

The NVIDIA Run:ai welcome screen shown after the first successful sign-in

Figure 10: The welcome screen confirms that the captured demo reached authenticated NVIDIA Run:ai onboarding. It does not show a completed project setup, a submitted workload, or GPU performance results.

Complete the onboarding appropriate to your environment, then check that the registered cluster appears in NVIDIA Run:ai and that its GPU resources are visible. At this point, we have moved from installing components to checking whether the platform can be used.

Making the Stack your own

The same pattern works when your platform needs a different storage default, another application, or an additional configuration task. Start with the dependencies: what must exist first, what indicates readiness, and what value does the next component need?

Certified templates are read-only and can be replaced by a Platform upgrade. Create a separate template for your changes. In Management > Stacks & Apps, open the certified Stack's CR view, copy its YAML, and choose Create Stack Template. Give it a new metadata.name, remove the vcluster.com/certified annotation and server-managed metadata, and remove status if present in the export.

For the CLI route, reconnect to the Platform management API before exporting:

vcluster platform connect management
kubectl get stacktemplates.management.loft.sh \
 run-ai-dedicated-control-plane -o yaml > my-runai-stack.yaml

Edit the exported file as described above, then apply it:

kubectl apply -f my-runai-stack.yaml

To use it through tenant creation, also copy the runai-tenant VirtualClusterTemplate and change the StackTemplate reference in its deploy.stacks entry to your new template. Expose or pass the parameters your environment needs, then create a fresh test tenant from that copied template. Updating a StackTemplate without connecting it to a StackInstance or tenant template does not deploy it anywhere.

When adding tasks, use dependsOn for actual readiness dependencies. Use declared parameters for values supplied by the user, and outputs for values discovered during deployment. If a task consumes another task's output, it must also declare that dependency. Keep credentials in the sensitive Secret-output path.

To note: StackTemplates are resolved during reconciliation and do not have Stack revisions. A change can affect existing instances. Test your copy with a fresh tenant, review pruning behavior, and retain the tested NVIDIA Run:ai chart versions and task names unless you have validated the resulting changes. Changing runaiVersion alone does not update the pinned charts. Your custom copy also becomes your responsibility to maintain when the certified bundle changes.

Step 6: Create a project and run something useful

Once NVIDIA Run:ai is available, create a project for the team or experiment that will use it. An NVIDIA Run:ai project maps workloads to a Kubernetes namespace and controls access and resource allocation. It is separate from the vCluster Platform project that contains the tenant cluster.

In NVIDIA Run:ai, open Organization > Projects, select the registered cluster, and choose +NEW PROJECT. For this example:

  • Name the project gpu-demo and choose the appropriate organizational scope.
  • Let NVIDIA Run:ai create its namespace, or follow the UI command to associate an existing one.
  • Continue to resource assignment and allocate at least one GPU from the intended node pool for the first workload.
  • Grant the intended user access and confirm that the project is ready.

Use the NVIDIA Run:ai project guide for the fields in your release. Creating a project without resource assignment can leave it with zero guaranteed GPU quota.

Final thoughts

The value of a Stack is not that it installs NVIDIA Run:ai. It is that the second tenant costs almost nothing. Dependency order, readiness gates, and the values passed between tasks are written down once, so adding the tenth team looks like adding the first: pick the template, supply the environment-specific inputs, watch the task graph. Start with one project and one workload, verify the result, and use that working configuration as the foundation for the next team.

Authoring your own is also a simple step. A StackTemplate is a list of tasks with dependsOn edges, a few declared parameters, and the outputs later tasks read. Start from the certified NVIDIA Run:ai template or the starter example, change one thing, test it on a fresh tenant cluster, and keep going. Your copy is a separate resource, so you own its upgrade path in exchange for controlling what it deploys.

Stuck, or want to compare notes with others building on this? Join us in the vCluster Slack.

Share:
Get started with the
#1 platform for AI infra.

Trusted by today’s fastest-growing AI cloud builders.

Ready to take vCluster for a spin?

Deploy your first virtual cluster today.