Tech Blog by vClusterPress and Media Resources

Introducing Stacks: ship AI environments as products, not projects

Oct 8, 2026
|
6
min Read
Introducing Stacks: ship AI environments as products, not projects

Raw bare metal runs about 14% gross margin after depreciation, and for most providers more than half of revenue comes from one or two customers (McKinsey). The way out is well understood: sell managed 

clusters and managed AI platforms instead of raw capacity. Same megawatt, roughly 2.5 times the revenue, and the difference is the software layer, not the hardware (company filings, Bernstein). SemiAnalysis puts a number on the quality gap too, with top-tier providers commanding roughly 10 to 15% more per GPU-hour than their direct competition (ClusterMAX 2.0).

So why doesn't every operator ship more managed services?

Because shipping one means building a full environment, not flipping on a feature. A managed NVIDIA Run:ai offering means registry credentials, an ingress controller, certificates, the GPU operator, a control plane, cluster registration, and monitoring wired to endpoints that did not exist until three steps ago, plus a person who knows the order to do it in. Multiply that by every tenant you onboard and every service you add, and GPUs stop being the constraint on your product line. The number of engineers who can assemble an environment correctly becomes the constraint instead.

That is the work Stacks removes.

Define the environment once, hand it out repeatedly

A Stack is a complete tenant environment described as a single reusable definition: which applications, in what order, what has to be healthy before the next thing starts, and which values the person deploying it is allowed to set.

Deploy it and you get a working environment. Deploy it a hundred times and you get a hundred identical working environments, each one owned, monitored, and removable on its own.

That changes three things that matter commercially.

A new service becomes something you ship, not something you staff. Define the environment once and every tenant after that is a deployment. The engineering cost moves into authoring the template, one time, instead of repeating the install for every customer. Your catalog stops being limited by how many people know how to build an environment correctly.

Onboarding stops being the bottleneck. The economics of managed services depend on serving many customers on short terms rather than a few on multi-year ones. That only works if the tenth customer costs the same to onboard as the first.

You keep control of what you handed out. The template author decides which parameters a consumer can set, and everything else stays fixed. A deployed environment stays a managed object with its own health, logs, and ownership, so removing a tenant removes exactly that tenant's resources and nothing else. Published values like endpoints and generated credentials sit behind their own permission, separate from the ability to see that an environment exists.

For enterprise platform teams the same mechanics show up as governance rather than margin. Every team gets the same environment, built the way your platform group decided it should be built, with the knobs you chose to expose and no others. Nobody edits Helm values to get a GPU environment, and nobody files a ticket to get one either.

Four ways to run NVIDIA Run:ai, ready to deploy

Certified Stacks are pre-validated Stack templates that vCluster tests and ships with the platform, so a supported integration ships as a deployment, without the integration project that used to come with it.

vCluster Platform 4.12 ships four NVIDIA Run:ai Certified Stacks, covering both ways operators actually run it:

  • Self-hosted, all in one. Everything NVIDIA Run:ai needs on a single cluster, with no manual preflight. Use it when one team, one cluster, and a working GPU platform is the whole requirement.
  • Shared host foundation. Install the NVIDIA Run:ai control plane and GPU Operator once, and let every tenant cluster share it. Use it when NVIDIA Run:ai is a service you sell to many customers.
  • Tenant registration. Connect a tenant cluster to that shared control plane, and disconnect only that tenant when they leave. Use it to make onboarding and offboarding a routine, repeatable, operation.
  • Tenant cluster components. The lightweight piece inside each tenant cluster, configured automatically from the shared foundation. Use it because nobody should be hand-copying values into a customer's environment.

Together they mean a provider can stand up managed NVIDIA Run:ai once and onboard the next customer in minutes, on hardware they already own.

Your product, deployed the same way everywhere

The more interesting part is that anyone can author a Stack, not just us.

If your software runs inside your customers' clusters, a Stack is where you encode what a correct installation looks like. Authoring one means declaring the tasks, the order they run in, and the parameters you want exposed. Ship that once and your product lands identically in every environment it reaches, whether the customer is your first or your hundredth. Your support burden drops, and so does the time between a customer deciding to use you and actually using you. Saturn Cloud is one: its AI token factory now deploys as a Stack, so an operator can give each customer private, per-token inference in an isolated tenant cluster.

Available Stacks

The NVIDIA Run:ai Certified Stacks are where this starts. The vCluster Stacks repository includes ready-made Stacks for additional environments, and examples from our own engineering team and  partners:

  • NVIDIA Dynamo, standing up a distributed inference environment inside a tenant cluster so an operator can offer serving for large models that span many GPUs and nodes.
  • Saturn Cloud, installing the AI token factory platform inside a tenant cluster so an operator can offer per-token inference, fine-tuning, and usage billing on GPUs they already own.
  • Community examples. The public repository is open for example Stack templates from anyone, whether you are building a Stack for your own product or standing up an environment your team keeps rebuilding.

If your product belongs in a tenant environment, it belongs in a Stack. Check out the vCluster Stacks repository, or talk to us about building one.

What you get, and what you still own

Stacks give you consistent provisioning and onboarding, plus an environment that stays inspectable afterward: aggregate health across every component, per-component logs, the ability to retry only the piece that failed, and clean removal.

Stacks are not a drift-reconciliation engine. If you want Git as the continuous source of truth for what runs inside a tenant cluster, put Argo CD applications inside your Stack and let Argo do what Argo is good at. Stacks and GitOps work together.

When something does go wrong, the troubleshooting runbook covers what each status means and where to look.

Get started

Stacks ship in vCluster Platform v4.12 with vCluster v0.37. There is no separate license. The Argo CD integration and Fleet Observability are now included in the Free tier, so registering clusters with Argo CD and monitoring a fleet no longer requires a paid license.

Deploy a Certified Stack, or talk to us about building a Stack for your own product.

Share:
Get started with the
#1 platform for AI infra.

Trusted by today’s fastest-growing AI cloud builders.

Ready to take vCluster for a spin?

Deploy your first virtual cluster today.