DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3 Provider Pricing Guide: Costs ComparedLatest article
Published on 2026.10.03 by DeepInfraGLM-5.3 Provider Pricing Guide: Costs Compared

GLM-5.3 is a large-scale reasoning model from Z.ai, released on August 18, 2026. Artificial Analysis tracks it as GLM-5.3 (max), emphasizing the reasoning variant, while OpenRouter lists it as Z.ai: GLM 5.3 and DeepInfra hosts it as zai-org/GLM-5.3. The model is open-weights, uses a Mixture of Experts architecture with 753 billion total parameters and 40 […]

Recent articles
GLM-5.3 Is Now Available on DeepInfraPublished on 2026.10.03 by DeepInfraGLM-5.3 Is Now Available on DeepInfra

GLM-5.3 shares its base model with GLM-5.2, so every performance gain comes from post-training. Z.ai scaled the reinforcement learning stack it already had and released the result on August 14, 2026. GLM-5.3 is a reasoning model built for complex software engineering and long-horizon agentic tasks. On Z.ai’s internal Code Bench, it improves 50% over GLM-5.2, […]

GLM-5.3 API Providers: Speed, Latency & CostPublished on 2026.10.02 by DeepInfraGLM-5.3 API Providers: Speed, Latency & Cost

GLM-5.3 API Review Summary GLM-5.3 is Z AI’s flagship coding and agentic reasoning model, released on August 14, 2026, and available on DeepInfra since launch. The model is built on the same ~753-billion parameter Mixture of Experts (MoE) base architecture as GLM-5.2, with all performance gains derived entirely from scaled post-training rather than architectural changes. […]

Best DeepSeek-V4.1-Flash API Providers in 2026Published on 2026.10.02 by DeepInfraBest DeepSeek-V4.1-Flash API Providers in 2026

As LLM architectures grow increasingly sophisticated, deploying state-of-the-art models like DeepSeek-V4.1-Flash requires more than just a basic API wrapper. For engineering teams, the challenge lies in balancing time-to-first-token (TTFT), throughput, context caching, and enterprise-grade compliance. DeepSeek-V4.1-Flash offers remarkable capabilities—including a massive 1M+ token context window and Engram conditional memory—but unlocking its full potential depends heavily […]

DeepSeek-V4.1-Flash: Model Overview & IntegrationPublished on 2026.10.01 by DeepInfraDeepSeek-V4.1-Flash: Model Overview & Integration

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone and a context window of 1,048,576 tokens. It accepts text and image input and generates text autoregressively. DeepInfra serves the model at fp8 with the full 1,048,576-token context, so the served window matches the model’s native window. The architecture targets input-heavy agentic workloads: the Causal […]

DeepSeek V4.1 Flash API: Speed, Latency & CostPublished on 2026.09.30 by DeepInfraDeepSeek V4.1 Flash API: Speed, Latency & Cost

DeepSeek V4.1 Flash (Reasoning, Max Effort) API Review Summary Metric Value Context Intelligence 40 (Artificial Analysis Intelligence Index) Well above median for comparable open-weight models (median: 18) Speed 211.5-545.6 output tokens/sec Notably fast; median: 68.9 t/s Latency (TTFT) 1.19s-5.28s (varies by provider) Competitive; median: 2.32s Price (DeepSeek API) $0.30/1M input, $1.20/1M output (peak) Cache discount: […]

GLM-5.3-Flash Documentation & Integration GuidePublished on 2026.09.30 by DeepInfraGLM-5.3-Flash Documentation & Integration Guide

GLM-5.3-Flash is a frontier-class, natively multimodal model developed by Z.ai and hosted on DeepInfra. It uses a Mixture-of-Experts (MoE) architecture with 320 billion total parameters, of which only 18 billion are active during inference. The model is built for complex, long-horizon tasks — advanced software engineering, agentic workflows, and multimodal reasoning — with a context […]