DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
The Qwen Team’s latest open-weight release, Qwen3.8-27B, is a 27-billion-parameter vision-language model. It goes deep on the tasks that trip up most models: multi-step agentic workflows, software engineering, and multimodal reasoning across images, documents, and video. On SWE-bench Pro, it scores 61.7, outpacing Opus4.6 Max’s 53.4. On QwenSWEBench it jumps from 49.3 on the previous […]
Published on 2026.10.03 by DeepInfraGLM-5.3 Provider Pricing Guide: Costs ComparedGLM-5.3 is a large-scale reasoning model from Z.ai, released on August 18, 2026. Artificial Analysis tracks it as GLM-5.3 (max), emphasizing the reasoning variant, while OpenRouter lists it as Z.ai: GLM 5.3 and DeepInfra hosts it as zai-org/GLM-5.3. The model is open-weights, uses a Mixture of Experts architecture with 753 billion total parameters and 40 […]
Published on 2026.10.03 by DeepInfraGLM-5.3 Is Now Available on DeepInfraGLM-5.3 shares its base model with GLM-5.2, so every performance gain comes from post-training. Z.ai scaled the reinforcement learning stack it already had and released the result on August 14, 2026. GLM-5.3 is a reasoning model built for complex software engineering and long-horizon agentic tasks. On Z.ai’s internal Code Bench, it improves 50% over GLM-5.2, […]
Published on 2026.10.02 by DeepInfraGLM-5.3 API Providers: Speed, Latency & CostGLM-5.3 API Review Summary GLM-5.3 is Z AI’s flagship coding and agentic reasoning model, released on August 14, 2026, and available on DeepInfra since launch. The model is built on the same ~753-billion parameter Mixture of Experts (MoE) base architecture as GLM-5.2, with all performance gains derived entirely from scaled post-training rather than architectural changes. […]
Published on 2026.10.02 by DeepInfraBest DeepSeek-V4.1-Flash API Providers in 2026As LLM architectures grow increasingly sophisticated, deploying state-of-the-art models like DeepSeek-V4.1-Flash requires more than just a basic API wrapper. For engineering teams, the challenge lies in balancing time-to-first-token (TTFT), throughput, context caching, and enterprise-grade compliance. DeepSeek-V4.1-Flash offers remarkable capabilities—including a massive 1M+ token context window and Engram conditional memory—but unlocking its full potential depends heavily […]
Published on 2026.10.01 by DeepInfraDeepSeek-V4.1-Flash: Model Overview & IntegrationDeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone and a context window of 1,048,576 tokens. It accepts text and image input and generates text autoregressively. DeepInfra serves the model at fp8 with the full 1,048,576-token context, so the served window matches the model’s native window. The architecture targets input-heavy agentic workloads: the Causal […]
Published on 2026.09.30 by DeepInfraDeepSeek V4.1 Flash API: Speed, Latency & CostDeepSeek V4.1 Flash (Reasoning, Max Effort) API Review Summary Metric Value Context Intelligence 40 (Artificial Analysis Intelligence Index) Well above median for comparable open-weight models (median: 18) Speed 211.5-545.6 output tokens/sec Notably fast; median: 68.9 t/s Latency (TTFT) 1.19s-5.28s (varies by provider) Competitive; median: 2.32s Price (DeepSeek API) $0.30/1M input, $1.20/1M output (peak) Cache discount: […]
© 2026 DeepInfra. All rights reserved.