DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3-Flash API Providers: Speed & CostPublished on 2026.09.29 by DeepInfraGLM-5.3-Flash API Providers: Speed & Cost

API Review Summary Metric Value Intelligence (Artificial Analysis Intelligence Index) 42 — well above the open-weight median (18) Speed 55.9 output tokens/sec — slower than the median (85.7 t/s) Latency (TTFT) 3.14s — higher than the median (2.05s) Cost (Z.ai first-party API) $0.15 / 1M input, $0.50 / 1M output; cache discount ~83% Cost efficiency […]

DeepSeek-V4.1-Flash Is Now on DeepInfraPublished on 2026.09.28 by DeepInfraDeepSeek-V4.1-Flash Is Now on DeepInfra

DeepSeek’s new V4.1-Flash uses just 8 billion active parameters during input processing, out of a 552-billion-parameter backbone, and still beats the much larger DeepSeek-V4-Pro across every agentic benchmark DeepSeek published. Released in September 2026, it is built around a Causal Encoder-Decoder design that makes the gap between total and active parameters possible. For developers running […]

GLM-5.3-Flash API Is Now on DeepInfraPublished on 2026.09.28 by DeepInfraGLM-5.3-Flash API Is Now on DeepInfra

Z.ai’s GLM-5.3-Flash activates just 18 billion of its 320 billion parameters at inference time, and on Z.ai’s reported results it outscores Claude Opus 4.8 on DeepSWE v1.1 (63.4 vs. 58.0), a demanding software engineering benchmark. It’s the first natively multimodal model in the GLM-5 series, combining text, image, and video input with a one-million-token context […]

Small Open-Weight Models: The Hidden Cost AdvantagePublished on 2026.09.28 by Stefan FidanovSmall Open-Weight Models: The Hidden Cost Advantage

Qwen3.8-27B landed at #9 on Arena.ai’s Code Arena WebDev board with 1,595 points, and as of September 2026 it’s the only model under 30 billion parameters in the top 10. Gemma 4-31B, released the same year and larger by three billion parameters, sits at #80. Three ranks above the 27B sits Qwen3.8-2.4T-A95B, a sibling with […]

Agents Need a Runtime Boundary, Not Just GuardrailsPublished on 2026.09.28 by DeepInfraAgents Need a Runtime Boundary, Not Just Guardrails

As agents move beyond model calls to tools, APIs, and execution environments, the runtime becomes part of the architecture. How DeepInfra Sandboxes and Hosted Agents keep the boundary outside the agent, where NVIDIA OpenShell fits, and why we plan around NVIDIA Vera.

AI Model Calibration: The Benchmark Nobody OptimizesPublished on 2026.09.25 by NiklasAI Model Calibration: The Benchmark Nobody Optimizes

DeepSeek V4 Pro scores 42 on the Artificial Analysis Intelligence Index. On the AA-Omniscience benchmark, which asks models hard factual questions and measures whether they answer or admit uncertainty, it has a 95% hallucination rate. That means when V4 Pro does not have the answer, it guesses anyway roughly 95 times out of 100. GPT-5.6 […]