Summary
- Raw GPU rental yields only 14-16% gross margin after depreciation, and GPU-hour prices are projected to fall 50% over five years. Bare metal alone becomes a shrinking-margin business.
- The winning providers move up the stack into managed clusters, Ray, inference, and training: higher-margin, stickier products built on top of GPU capacity.
- Production Ray-as-a-service requires hard tenant isolation: each customer gets a dedicated Ray cluster with its own API server, scheduler, RBAC, and worker nodes, with no cross-tenant scheduling.
- vCluster Platform delivers that isolation with Private Nodes as the production default, so AI cloud providers can launch managed Ray on their existing GPU estate without building the isolation layer themselves.
Raw GPU rental generates 14-16% gross margin after depreciation. GPU-hour prices are projected to fall 50% over the next five years. If selling bare metal compute is your entire business, the math gets worse every quarter.
The providers gaining ground are moving up the stack. Bare metal is the foundation. Managed clusters, managed Ray, managed inference, and managed training are the products, each one a higher-margin, stickier offering than the GPU-hour beneath it. Your customers running distributed ML workloads are already asking for managed Ray. The question is whether you build it correctly or hand that revenue to someone else.
This article is written for AI cloud providers who own or operate GPU racks and want to sell Ray-as-a-service to their customers, not for the teams consuming Ray. The architecture, the isolation requirements, and the go-to-market path are all framed from the supplier's side.
What "Ray-as-a-Service" Actually Means When You Are the Provider
From the end-user perspective, Ray is an open-source distributed compute framework that handles distributed training (Ray Train), data processing (Ray Data), and model serving (Ray Serve) across a cluster of workers. Customers want it managed so they can focus on their workloads rather than on cluster operations.
From your perspective as the provider, selling managed Ray means something more specific: you give each customer their own isolated Ray cluster (head node, worker nodes, scheduler, namespaces, and RBAC) with dedicated worker nodes per tenant. The customer experience is a private Ray cluster with full isolation. That isolation is where the product engineering lives.
The default path most teams attempt starts with the KubeRay operator on Kubernetes. The operator installs once, cluster-scoped, in an admin namespace. Each tenant gets their own namespace and can create RayCluster resources within it. The operator manages all of them from a single process. That setup is a starting point, and the finished product takes more work.
The Isolation Problem Providers Underestimate
A single cluster-scoped KubeRay operator managing every tenant's RayCluster resources does not meet the isolation bar for production tenant-isolated services. One operator process has visibility into all cluster resources it manages. Namespace boundaries alone do not provide the per-tenant API server, etcd, and control plane isolation that paying customers require. A misconfigured RBAC policy, an overly permissive service account, or a CNI-level operator deployed by one tenant can expose internals to another.
This is a well-understood problem in Kubernetes operations. When teams deploy service meshes or CNI-level operators at cluster scope, cross-tenant visibility becomes a genuine operational risk. For an internal environment with trusted teams, it is a manageable trade-off. For a commercial product serving paying customers running production training and inference workloads, customers who have no relationship with each other, it is not manageable.
Customers running production Ray workloads on your infrastructure need hard guarantees: their scheduler is theirs, their RBAC controls only their resources, and their worker nodes are dedicated to their workloads with no cross-tenant co-location. Without those guarantees, you cannot make a meaningful SLA commitment, and sophisticated customers will not move production workloads onto the platform.
What DIY Isolation Actually Costs
Building hard tenant isolation from scratch on top of raw KubeRay is the path many providers attempt first. The toolchain assembles quickly on a whiteboard: Capsule or Kyverno for tenant policy, custom Terraform and GitOps pipelines for per-tenant RBAC, separate operator instances or namespace-scoped deployments for the KubeRay operator, and custom networking policy to enforce data-plane separation.
The build phase is just the beginning; every new cluster type, every new customer security requirement, and every upstream upgrade adds to the maintenance load indefinitely. In concrete terms, building a production-grade, isolated cluster platform takes 6-10 platform engineers and 6-12 months to reach a launchable state — and at realistic scope, over $1M in engineering before accounting for ongoing maintenance. While that layer is under construction, a $10M GPU cluster keeps earning commodity GPU-hour margin instead of the higher-margin managed-service revenue it could be producing.
How vCluster Turns Ray into a Managed Product
vCluster Platform is purpose-built for AI cloud providers who need to productize cluster experiences (Kubernetes, Slurm (Beta), Run:AI, Ray, and more) on their own hardware. It handles the isolation layer that the DIY path requires providers to build themselves.
The core mechanism is control plane virtualization. Instead of giving each customer a separate physical cluster or running multiple tenants through a single control plane, vCluster gives each tenant cluster a fully isolated control plane: its own API server, its own scheduler, its own data store, and its own RBAC, with no separate physical clusters required. From the tenant's perspective, they have a dedicated cluster. From the provider's perspective, each tenant cluster runs on dedicated worker nodes (Private Nodes), delivering full isolation.
Ray is a GA cluster type in vCluster via Certified Stacks, pre-validated integrations that deploy production-ready environments with no manual setup and no configuration drift. Other Certified Stacks include Run:AI, Jupyter, and Slurm (Beta) via Slinky. To be precise about the architecture: vCluster provides the isolation envelope around the Ray cluster. Ray's distributed scheduling and execution capabilities are Ray's own, developed by Anyscale. vCluster does not replace those capabilities; it secures them for a commercial, multi-customer offering by placing each customer's Ray cluster inside a fully isolated tenant cluster.
Private Nodes: The Production Default
vCluster's production default is hard isolation. That means Private Nodes: dedicated worker nodes per tenant, with a per-tenant CNI and per-tenant storage. Private Nodes delivers hardware-level isolation that approaches a dedicated physical cluster: dedicated worker nodes joined directly and privately into each tenant cluster, not visible from the control plane cluster, with per-tenant CNI and storage, and no cross-tenant scheduling. Each customer's Ray workers run on hardware that no other tenant's processes touch.
For customers running production training workloads or latency-sensitive inference, this is the configuration that supports a real SLA.
What Each Tenant Cluster Looks Like
When a customer purchases managed Ray from you, they receive:
- A dedicated Ray head node and worker nodes running inside their own isolated tenant cluster
- Their own Kubernetes API server and scheduler, invisible to other tenants
- Their own RBAC, so they can create service accounts, roles, and bindings without affecting any other customer
- Their own CRD scope, so their
RayClusterresources do not interact with another tenant's - Per-tenant CNI and storage, enforced at the node level via Private Nodes
Tenant clusters can be provisioned via CI/CD, APIs, or a self-service portal in seconds. The Certified Stack for Ray means every customer receives the same validated environment, policies, and configuration, with no manual setup and no per-customer drift.
The Ray-as-a-Service Kubernetes Path Without the Plumbing
The challenge with building Ray-as-a-Service Kubernetes infrastructure from scratch is that the Kubernetes primitives (namespaces, RBAC, network policies) were not designed to provide the level of isolation that paying customers require between untrusted workloads. vCluster's control plane virtualization layer sits above those primitives and delivers per-tenant isolation that namespace separation alone cannot provide. Providers get the operational efficiency of managing many isolated tenant clusters on their GPU estate with the isolation guarantees of per-cluster deployment, without building or maintaining the isolation layer themselves.
Proof Points: Providers Who Have Already Made the Leap
AI cloud providers are already shipping managed cluster products on vCluster's isolation layer:
- Boost Run launched a GPU-native managed Kubernetes service in under 45 days with zero new platform engineering hires.
- QumulusAI spins up isolated Kubernetes environments for AI customers in 1 minute.
- Lintasarta launched Indonesia's leading AI cloud in 90 days and now runs 170+ tenant clusters in production.
The same isolation and orchestration layer that ships managed Kubernetes is the exact layer that ships managed Ray. vCluster Platform is production-proven at 100K+ GPUs across 50+ GPU Clouds and Fortune 500s, with 40M+ tenant clusters created.
The Revenue Ladder: Ray Is One Step, Not the Ceiling
Ray-as-a-service sits on a progression where each step unlocks higher margin than the one before:
- Bare metal: GPU hours at commodity rates, 14-16% gross margin
- Managed clusters: Kubernetes or Ray clusters with SLA commitments, sold at a premium over raw compute
- Isolate and pack: dense multi-tenant deployment with hard isolation — more paying tenants per GPU, lifting effective margin per unit of silicon
- Managed inference: GPU capacity sold as inference tokens, priced on output value rather than silicon cost
- Managed training: long-running, checkpointed training jobs sold as a managed service
vCluster Platform supports every layer of this progression, delivered as an integrated stack in top-down order. Certified Stacks are ready-to-run AI/ML environments (Run:AI, Ray, Jupyter, Slurm (Beta) via Slinky) and deploy in minutes. vNode provides kernel-native workload isolation via seccomp, cgroups, and namespaces with no hypervisor overhead, preserving bare metal GPU performance. vCluster handles tenant cluster orchestration for any cluster type. vMetal is the machine layer: zero-touch PXE boot, OS install, machine registration, and network automation via a single EC2-like API, producing bare metal machines and VMs as sellable products.
The full stack (Certified Stacks -> vNode -> vCluster -> vMetal) means providers are not assembling a platform from unrelated open-source components. It is an integrated infrastructure layer with a single management plane across public cloud, AI cloud, and on-premises GPU infrastructure.
The pricing model shift matters as much as the infrastructure shift. A GPU rented by the hour is priced based on the cost of the silicon. A GPU monetized by the token is priced based on the value of the output. Moving up the stack, starting with managed Ray, is what makes that pricing shift possible. Anyscale's own pricing model, offering usage-based billing with committed contracts for volume, is the end-user benchmark your customers will reference. Providers who can offer a comparable experience on their own hardware, at competitive rates and with strong isolation guarantees, capture that pricing power directly.
What to Do Next
AI cloud providers who are still selling only raw GPU hours are competing on a number that trends downward. The providers building managed services (starting with managed clusters, extending to managed Ray) are selling on capability, reliability, and isolation rather than on price.
The technical barrier is real: delivering hard tenant isolation on GPU infrastructure is not a problem that namespace-level Kubernetes separation solves. But it is a solved problem. vCluster Platform's Private Nodes, Certified Stacks, and control plane virtualization deliver what the DIY path requires providers to build from scratch.
If your customers are asking for managed Ray and your current answer is a KubeRay install that relies on namespace separation, the gap between what you are offering and what they need is the gap your competitors will fill.
Schedule an enterprise demo with the vCluster Labs team to see how Ray-as-a-service deploys on your existing GPU infrastructure, with the tenant isolation required to sell it as a production product.
Frequently Asked Questions
What is Ray-as-a-service for AI cloud providers?
Ray-as-a-service for AI cloud providers means selling each customer a managed, isolated Ray cluster (head node, worker nodes, scheduler, and RBAC) as a product. The provider owns and operates the GPU infrastructure, while the customer consumes Ray without managing cluster operations. In vCluster terms, this is delivered as a dedicated tenant cluster with Private Nodes as the production default isolation model.
Why is namespace-level isolation not enough for managed Ray?
Namespace boundaries alone do not provide the per-tenant API server, etcd, and control plane isolation that paying customers require. A single cluster-scoped KubeRay operator has visibility into every RayCluster across namespaces, creating cross-tenant visibility risk. Production Ray-as-a-service requires tenant-isolated environments with dedicated worker nodes and no cross-tenant scheduling.
How does vCluster provide tenant isolation for Ray?
vCluster runs on a control plane cluster and gives each customer a fully isolated tenant cluster with its own Kubernetes API server, scheduler, data store, RBAC, and CRD scope. The production default is Private Nodes, which dedicates worker nodes and per-tenant CNI and storage to a single tenant, eliminating cross-tenant scheduling.
What is Private Nodes in vCluster?
Private Nodes is vCluster's production default isolation model for AI cloud and GPU cloud providers. It joins dedicated worker nodes directly and privately into each tenant cluster, with per-tenant CNI and per-tenant storage. No other tenant's processes touch those nodes, giving each customer hardware-level isolation that supports meaningful SLAs.
Does vCluster replace Ray's scheduler or execution engine?
No. vCluster does not replace Ray's distributed scheduling, Ray Train, Ray Data, or Ray Serve capabilities. It provides the isolation envelope around each Ray cluster by placing it inside a fully isolated tenant cluster. Ray remains Ray; vCluster secures it for a commercial, multi-customer offering.
Can vCluster be used for services other than Ray?
Yes. vCluster Platform supports multiple Certified Stacks, including Run:AI, Jupyter, and Slurm (Beta) via Slinky, in addition to Ray. It provides the same tenant cluster orchestration and Private Nodes isolation model for Kubernetes, Slurm (Beta), Run:AI, and other AI cloud services.
How fast can a GPU cloud provider launch managed Ray with vCluster?
Deployments are accelerated because the Ray Certified Stack is pre-validated and Private Nodes is the production default. You avoid building a DIY isolation layer and can start selling managed Ray on existing GPU infrastructure quickly.
What pricing shift should AI cloud providers make with managed Ray?
Managed Ray lets providers move from GPU-hour pricing based on silicon cost to usage-based or committed-contract pricing based on output value. That shift is what turns a 14-16% gross margin raw GPU rental business into a higher-margin managed service selling capability, reliability, and tenant isolation.
Deploy your first virtual cluster today.