We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepSeek’s $10.29B Financing Round ExplainedPublished on 2026.07.01 by DeepInfraDeepSeek’s $10.29B Financing Round Explained

DeepSeek has not taken outside money since it was founded in 2023. For two years it turned down every venture capital firm and major tech company that came calling, funding its research entirely from the returns of its parent hedge fund, Zhejiang High-Flyer Asset Management, which reportedly posted a 56.6% return in 2025. That era […]

How DeepInfra Built on NVIDIA's Inference Stack and Why It Paid OffPublished on 2026.06.30 by Aray SultanbekovaHow DeepInfra Built on NVIDIA's Inference Stack and Why It Paid Off

When we built DeepInfra, we made a deliberate bet on the NVIDIA inference software stack. Not as a hedge — as a conviction. Today, that bet is paying off in ways that are easy to measure.

Introducing the Priority Service Tier: Front-of-Queue Inference When It CountsPublished on 2026.06.29 by DeepInfraIntroducing the Priority Service Tier: Front-of-Queue Inference When It Counts

Pay 1.5× real-time for priority scheduling and protected capacity.

Introducing the Batch API: Run Large Inference Jobs 20% CheaperPublished on 2026.06.19 by Vasilije NovakovicIntroducing the Batch API: Run Large Inference Jobs 20% Cheaper

DeepInfra's new Batch API lets you submit large volumes of completions, chat, and embedding requests as a single asynchronous job—processed within 24 hours at 20% off real-time pricing. It's fully OpenAI-compatible, so if you've used OpenAI's Batch API, you already know how it works.

Step 3.7 Flash is Live on DeepInfra: An Agentic, Multimodal Model Built for ProductionPublished on 2026.06.12 by DeepInfraStep 3.7 Flash is Live on DeepInfra: An Agentic, Multimodal Model Built for Production

StepFun's Step 3.7 Flash is now live on DeepInfra. It's a 198B-parameter sparse MoE vision-language model with just ~11B active parameters per token, a 256K context window, and three selectable reasoning levels—purpose-built for high-throughput agentic workflows that combine perception, search, and reasoning.

DeepInfra Launches Access to NVIDIA Cosmos 3 World Foundation Models for Physical AIPublished on 2026.06.04 by Yessen KanapinDeepInfra Launches Access to NVIDIA Cosmos 3 World Foundation Models for Physical AI

DeepInfra is serving NVIDIA Cosmos 3, the first open world foundation model for physical AI that reasons before it generates, from day zero of its release. Available as two variants—Cosmos 3 Nano and Cosmos 3 Super—these models give developers a cost-efficient foundation for building robots, autonomous vehicles, simulation workflows, and synthetic data generation at scale.