The integration pairs NVIDIA Run:ai’s GPU orchestration with Saturn Cloud’s multi-tenant inference platform and the NVIDIA DSX AI Factory Platform, giving neocloud and AI factory operators a way to earn more from their GPU fleets by turning raw capacity into per-token products they run under their own brand.
NEW YORK, Sept. 17, 2026 /PRNewswire/ — Saturn Cloud, the AI token factory platform, today announced an integration with NVIDIA Run:ai, the AI workload and GPU orchestration platform. The integration helps GPU cloud operators earn more from their installed capacity by turning fleets that would otherwise be rented by the GPU-hour into multi-tenant inference products sold by the token. NVIDIA Run:ai is one of several NVIDIA technologies Saturn Cloud builds on, following its integration of NVIDIA DSX AI Factory Platform.
NVIDIA Run:ai maximizes GPU utilization across the fleet through NVIDIA KAI Scheduler, providing gang scheduling for NVIDIA Grove-managed distributed workloads and policy-driven governance. Saturn Cloud runs on top, giving operators the commercial layer between their fleet and customers so they can sell inference without building it themselves.
Beyond per-token serving, operators get multiple products to sell from the same infrastructure — GPU hours, per-token inference endpoints, and managed fine-tuning jobs — without building any of it in-house. They can offer dedicated GPU capacity to customers who bring their own stack, per-token model-as-a-service to customers who just want an API endpoint, and self-service development environments for teams that build and fine-tune their own models. Saturn Cloud also handles model onboarding, shared and dedicated isolation tiers for regulated customers, and the identity, access, and governance controls enterprise buyers require.
Underneath, the serving layer runs on NVIDIA Dynamo, handling distributed inference with disaggregated prefill and decode, and drives vLLM, SGLang, or NVIDIA TensorRT-LLM. NVIDIA Grove orchestrates the multi-node workloads, and NVIDIA KAI scheduler places them with GPU awareness, allocates fractional GPUs, and enforces quotas across tenants. NVSentinel and NVIDIA Fleet Intelligence handle fleet-wide health monitoring and fault remediation, so degraded hardware is automatically pulled from placement.
“A neocloud can raise revenue per megawatt without adding a single GPU, just by changing how the capacity is sold,” said Sebastian Metti, Founder of Saturn Cloud. “Saturn Cloud gives operators the full commercial platform to do it, handling serving, multi-tenancy, and billing, all running under their own brand on NVIDIA Run:ai.”
“Cloud providers are looking for new ways to turn GPU infrastructure into differentiated AI services,” said Omri Geller, NVIDIA VP DSX OS Platform Software. “Together, NVIDIA Run:ai and Saturn Cloud help operators improve GPU utilization while delivering scalable, multi-tenant inference services under their own brand.”
The NVIDIA Run:ai-integrated Saturn Cloud platform is available now. Operators can learn more by visiting saturncloud.io.
About Saturn Cloud
Saturn Cloud is the AI token factory platform for AI clouds and enterprises. It turns GPU infrastructure into managed services for fine-tuning, model serving, and per-token billing with built-in enterprise security across public, private, and on-premises environments. Saturn Cloud is built on NVIDIA accelerated computing, including NVIDIA Run:ai and the NVIDIA DSX OS toolkit. Learn more at saturncloud.io.
View original content:https://www.prnewswire.com/news-releases/saturn-cloud-integrates-nvidia-runai-to-turn-nvidia-gpu-fleets-into-inference-businesses-302882589.html
SOURCE Saturn Cloud
