We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepInfra Launches Access to NVIDIA Cosmos 3 World Foundation Models for Physical AI
Published on 2026.06.04 by Yessen Kanapin
DeepInfra Launches Access to NVIDIA Cosmos 3 World Foundation Models for Physical AI

DeepInfra is serving NVIDIA Cosmos 3, NVIDIA's open world foundation model for physical AI, from day zero of its release. As the first omnimodel for physical AI that reasons before it generates, Cosmos 3 is live on DeepInfra today as two variants—Cosmos 3 Nano and Cosmos 3 Super—at the industry's best prices, empowering developers to build physical AI systems without compromising on budget or performance.

What Makes Cosmos 3 Different

Most generative models just generate. Cosmos 3 does something different: it reasons first, then generates. That distinction matters a great deal if you're building physical AI systems like robots or autonomous vehicles, where generating plausible-but-wrong outputs isn't just a quality issue—it's a safety one. As NVIDIA describes it, Cosmos 3 is the first OmniModel that unifies reasoning, world, and action generation in a single architecture.

Under the hood it uses a Mixture-of-Transformer architecture that combines an autoregressive reasoner with a diffusion-based generator. Inputs and outputs span text, image, video, audio, and action, making Cosmos 3 genuinely multimodal in both directions—not just for perception, but for generation and decision-making as well.

What It's Good At

Synthetic Video Data

Ranked #1 open world generation model for synthetic data generation. Use it to generate training data for physical AI at scale, without expensive real-world data collection.

Policy Backbone

Ranked #1 backbone for world action models. A strong foundation for robotics, embodied AI, and AV policy training.

Visual Reasoning

Ranked #1 open model for visual understanding on fixed infrastructure cameras—useful for smart city, warehouse, logistics deployments, infrastructure monitoring, and industrial automation.

Simulated Environments

Designed for closed-loop learning and simulation workflows. Pairs with NVIDIA AV Sim and Isaac Sim for training, testing, and evaluating physical AI systems in simulated environments before deployment.

Two Variants

Cosmos 3 Nano

The lighter variant. A good starting point for experimentation, fine-tuning, and latency-sensitive workloads.

Cosmos 3 Super

The full-capability variant. Tops the PAI Bench and R-Bench leaderboards. Use it where quality and reasoning performance are the priority.

Both are available on DeepInfra today via our standard API—the same setup as any other model, with no special configuration needed to get started.

Getting Started with NVIDIA Cosmos 3 on DeepInfra

Cosmos 3 Nano and Cosmos 3 Super are live on DeepInfra now. If you're building physical AI, robots, or AV systems and want to experiment with world modeling, reasoning, action generation, and synthetic data creation, this is a strong place to start.

Visit our models page to explore competitive rates for Cosmos 3 inference, or check out the DeepInfra docs to learn more about our complete model ecosystem and developer resources.

Related articles
NVIDIA Nemotron 3 Super 120B API Benchmarks: Latency & CostNVIDIA Nemotron 3 Super 120B API Benchmarks: Latency & Cost<p>About NVIDIA Nemotron 3 Super 120B A12B NVIDIA&#8217;s Nemotron 3 Super 120B A12B is an open-weight large language model released on March 11, 2026. It features 120B total parameters with only 12B active per forward pass, delivering exceptional compute efficiency for complex multi-agent applications such as software development and cybersecurity triaging. The model uses a [&hellip;]</p>
Kimi K2.6 is Now Available on DeepInfraKimi K2.6 is Now Available on DeepInfra<p>Kimi K2.6 can coordinate up to 300 sub-agents executing 4,000 steps in a single autonomous run — Moonshot AI&#8217;s answer to the gap between what frontier models can do in a chat window and what production agentic systems actually need. Built for long-horizon coding, deep research, and complex orchestration, the model is open source under [&hellip;]</p>
Kimi K2.5 API Benchmarks: Latency, Throughput & CostKimi K2.5 API Benchmarks: Latency, Throughput & Cost<p>About Kimi K2.5 Kimi K2.5 is Moonshot AI&#8217;s flagship open-source reasoning model, released in January 2026. It is a native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens. The model features a Mixture-of-Experts (MoE) architecture with 1 trillion total parameters and 32 billion activated parameters. Kimi K2.5 [&hellip;]</p>