DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Small Open-Weight Models: The Hidden Cost AdvantagePublished on 2026.09.28 by Stefan FidanovSmall Open-Weight Models: The Hidden Cost Advantage

Qwen3.8-27B landed at #9 on Arena.ai’s Code Arena WebDev board with 1,595 points, and as of September 2026 it’s the only model under 30 billion parameters in the top 10. Gemma 4-31B, released the same year and larger by three billion parameters, sits at #80. Three ranks above the 27B sits Qwen3.8-2.4T-A95B, a sibling with […]

Agents Need a Runtime Boundary, Not Just GuardrailsPublished on 2026.09.28 by DeepInfraAgents Need a Runtime Boundary, Not Just Guardrails

As agents move beyond model calls to tools, APIs, and execution environments, the runtime becomes part of the architecture. How DeepInfra Sandboxes and Hosted Agents keep the boundary outside the agent, where NVIDIA OpenShell fits, and why we plan around NVIDIA Vera.

AI Model Calibration: The Benchmark Nobody OptimizesPublished on 2026.09.25 by NiklasAI Model Calibration: The Benchmark Nobody Optimizes

DeepSeek V4 Pro scores 42 on the Artificial Analysis Intelligence Index. On the AA-Omniscience benchmark, which asks models hard factual questions and measures whether they answer or admit uncertainty, it has a 95% hallucination rate. That means when V4 Pro does not have the answer, it guesses anyway roughly 95 times out of 100. GPT-5.6 […]

Token Verbosity Is the New Pricing WarPublished on 2026.09.25 by DeepInfraToken Verbosity Is the New Pricing War

With each new model release, we see new benchmark claims and new arguments justifying the value of its tokens. However, when teams compare models for real-world workloads, the sticker price alone hides large per-task differences in the amount of work a model requires to solve a task. Cheaper tokens do not necessarily translate to lower […]

GLM-5.3-Flash Pricing, Providers & CostPublished on 2026.09.24 by DeepInfraGLM-5.3-Flash Pricing, Providers & Cost

GLM-5.3-Flash is a model from Z.ai in the GLM-5 family, released on August 26, 2026. It’s a natively multimodal model accepting text and image input (some listings, including OpenRouter, also note video), with text output. Both Artificial Analysis and DeepInfra describe it as a 320B-parameter mixture-of-experts model with 18B active parameters per inference; DeepInfra’s listing […]

Model Deprecation: Build LLM Apps That LastPublished on 2026.09.24 by Stefan FidanovModel Deprecation: Build LLM Apps That Last

Your model ID is the shortest-lived dependency in your stack and odds are it doesn’t have a maintenance schedule. On June 15, 2026, claude-sonnet-4-20250514 and claude-opus-4-20250514 stopped answering requests. Anthropic had posted the notice 62 days earlier. Teams with either string in a call site learned about it from an error rate, not an email. […]