DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepSeek V4.1 Flash API: Speed, Latency & CostLatest article
Published on 2026.09.30 by DeepInfraDeepSeek V4.1 Flash API: Speed, Latency & Cost

DeepSeek V4.1 Flash (Reasoning, Max Effort) API Review Summary Metric Value Context Intelligence 40 (Artificial Analysis Intelligence Index) Well above median for comparable open-weight models (median: 18) Speed 211.5-545.6 output tokens/sec Notably fast; median: 68.9 t/s Latency (TTFT) 1.19s-5.28s (varies by provider) Competitive; median: 2.32s Price (DeepSeek API) $0.30/1M input, $1.20/1M output (peak) Cache discount: […]

Recent articles
GLM-5.3-Flash Documentation & Integration GuidePublished on 2026.09.30 by DeepInfraGLM-5.3-Flash Documentation & Integration Guide

GLM-5.3-Flash is a frontier-class, natively multimodal model developed by Z.ai and hosted on DeepInfra. It uses a Mixture-of-Experts (MoE) architecture with 320 billion total parameters, of which only 18 billion are active during inference. The model is built for complex, long-horizon tasks — advanced software engineering, agentic workflows, and multimodal reasoning — with a context […]

DeepSeek-V4.1-Flash Pricing Guide for DevelopersPublished on 2026.09.29 by DeepInfraDeepSeek-V4.1-Flash Pricing Guide for Developers

If you’re evaluating long-context reasoning models in late 2026, DeepSeek-V4.1-Flash is hard to ignore because the pricing is aggressive, the weights are open, and the provider market around it is already competitive. This is the rare model that shows up in both cost-sensitive buying conversations and serious agentic workloads: posted API pricing starts as low […]

Best GLM-5.3-Flash API Providers in 2026Published on 2026.09.29 by DeepInfraBest GLM-5.3-Flash API Providers in 2026

As AI architectures shift toward highly optimized Mixture-of-Experts (MoE) models, GLM-5.3-Flash has emerged as a strong option for developers who need fast inference, advanced reasoning, and robust tool-calling. Deploying a model like this in production still means balancing token costs, time-to-first-token (TTFT) latency, raw throughput, and API reliability. The inference-cloud and API-gateway ecosystem for this […]

GLM-5.3-Flash API Providers: Speed & CostPublished on 2026.09.29 by DeepInfraGLM-5.3-Flash API Providers: Speed & Cost

API Review Summary Metric Value Intelligence (Artificial Analysis Intelligence Index) 42 — well above the open-weight median (18) Speed 55.9 output tokens/sec — slower than the median (85.7 t/s) Latency (TTFT) 3.14s — higher than the median (2.05s) Cost (Z.ai first-party API) $0.15 / 1M input, $0.50 / 1M output; cache discount ~83% Cost efficiency […]

DeepSeek-V4.1-Flash Is Now on DeepInfraPublished on 2026.09.28 by DeepInfraDeepSeek-V4.1-Flash Is Now on DeepInfra

DeepSeek’s new V4.1-Flash uses just 8 billion active parameters during input processing, out of a 552-billion-parameter backbone, and still beats the much larger DeepSeek-V4-Pro across every agentic benchmark DeepSeek published. Released in September 2026, it is built around a Causal Encoder-Decoder design that makes the gap between total and active parameters possible. For developers running […]

GLM-5.3-Flash API Is Now on DeepInfraPublished on 2026.09.28 by DeepInfraGLM-5.3-Flash API Is Now on DeepInfra

Z.ai’s GLM-5.3-Flash activates just 18 billion of its 320 billion parameters at inference time, and on Z.ai’s reported results it outscores Claude Opus 4.8 on DeepSWE v1.1 (63.4 vs. 58.0), a demanding software engineering benchmark. It’s the first natively multimodal model in the GLM-5 series, combining text, image, and video input with a one-million-token context […]