DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Multi-Turn RL: A Guide to Getting Reinforcement Learning RightLatest article
Published on 2026.09.08 by Stefan FidanovMulti-Turn RL: A Guide to Getting Reinforcement Learning Right

The first multi-turn RL run you launch will spend most of its life doing something you would not call training. You watch GPU utilization sit under half, watch a step take eleven minutes, and go hunting for a bug in your gradient accumulation. There is no bug. The trainer is waiting on rollouts. This is […]

Recent articles
Best AI Inference Platforms for Speed & Cost in 2026Published on 2026.09.08 by Stefan FidanovBest AI Inference Platforms for Speed & Cost in 2026

Your monitoring dashboard says the API is fast and your invoice says the same thing in different units. Somewhere between the two, cost per completed request disappears. This guide compares the best AI inference platform options for speed and cost on the same footing: measured latency, measured throughput, current prices, and the math that turns […]

Sandboxes: give your agents a safe place to run codePublished on 2026.08.19 by DeepInfraSandboxes: give your agents a safe place to run code

Isolated Linux microVMs you can spin up with one API call: run bash or Python, move files in and out, pause and resume, tear down when you're done. Built for agents and pipelines that need to execute untrusted code without you having to run the infrastructure yourself.

Best Kimi K3 SaaS Tools & API PlatformsPublished on 2026.08.13 by DeepInfraBest Kimi K3 SaaS Tools & API Platforms

Kimi K3, with its 2.8 trillion parameters and 1M-token context window, represents a significant leap in large language model capabilities. However, deploying, accessing, and managing a model of this scale presents real infrastructure challenges. From optimizing inference latency and managing GPU compute costs to handling multimodal vision capabilities, selecting the right deployment platform is critical […]

Kimi K3 Pricing, Providers & Real-World CostsPublished on 2026.08.12 by DeepInfraKimi K3 Pricing, Providers & Real-World Costs

Kimi K3 matters because it pushes an unusual combination into the same decision: open weights, a 1 million token context window, and frontier-class benchmark numbers, but at pricing still high enough to force real provider shopping. Released by Moonshot AI on July 16, 2026, it is a 2.8 trillion parameter Mixture-of-Experts model with 104 billion […]

Kimi K3: 2.8T Open-Weight Multimodal ModelPublished on 2026.08.11 by DeepInfraKimi K3: 2.8T Open-Weight Multimodal Model

Kimi K3, developed by Moonshot AI, represents a landmark achievement in open-source artificial intelligence. As a 2.8-trillion-parameter native multimodal Mixture-of-Experts (MoE) model, Kimi K3 is engineered to handle demanding computational tasks, from complex software engineering and long-horizon agentic workflows to deep scientific research. By combining a one-million-token context window with its architectural innovations, Kimi K3 […]

NVIDIA Nemotron 3.5 Lightning Is Live on DeepInfra: Day-Zero Access to the Fastest Open Model for AgentsPublished on 2026.08.11 by Aray SultanbekovaNVIDIA Nemotron 3.5 Lightning Is Live on DeepInfra: Day-Zero Access to the Fastest Open Model for Agents

NVIDIA Nemotron 3.5 Lightning is available on DeepInfra's serverless API from day zero. It's a 30B hybrid MoE model with 3B active parameters, a 1M-token context window, and up to 4x higher throughput for high-volume agentic workloads.