DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3-Flash is a frontier-class, natively multimodal model developed by Z.ai and hosted on DeepInfra. It uses a Mixture-of-Experts (MoE) architecture with 320 billion total parameters, of which only 18 billion are active during inference.
The model is built for complex, long-horizon tasks — advanced software engineering, agentic workflows, and multimodal reasoning — with a context window of just over one million tokens (1,048,576) and a serving stack tuned for efficiency. Z.ai positions it as delivering strong intelligence at a fraction of the cost of GLM-5.2, its immediate predecessor.
GLM-5.3-Flash introduces several changes over GLM-5.2 aimed at cutting compute without giving up reasoning quality.
On DeepInfra, GLM-5.3-Flash is available with private endpoint deployment, JSON mode, and function/tool calling.
Z.ai positions GLM-5.3-Flash as a “Flash” model that competes with frontier-class systems — Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash — on coding and agentic workflows, at a fraction of the parameter count active per token.
The table below reflects Z.ai’s own published benchmark results for GLM-5.3-Flash.
| Benchmark Category | GLM-5.3-Flash | GLM-5.2 | Claude Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash |
| Terminal-Bench 2.1 | 84.3 | 81.0 | 85.0 | 87.4 | 85.8 |
| DeepSWE v1.1 | 63.4 | 46.2 | 58.0 | 69.6 | 65.3 |
| Toolathlon (Verified) | 78.4 | 59.9 | 76.2 | 74.9 | – |
| AutomationBench v1.0.6 | 48.8 | 26.2 | 41.0 | 37.2 | 52.3 |
| MMVU (Vision) | 80.5 | – | 67.4 | 75.8 | 82.3 |
| CharXiv Reasoning (w/ Tools) | 89.4 | – | – | – | – |
| Chartography (w/ Tools) | 78.0 | – | – | – | – |
Note: an earlier version of this table included a row labeled “OSWorld 2.0” with scores for each model. That benchmark does not appear anywhere in Z.ai’s official results for GLM-5.3-Flash, and the figures could not be verified against any source — it’s been removed and replaced with CharXiv Reasoning and Chartography, two vision benchmarks Z.ai does report.
DeepInfra provides an OpenAI-compatible interface for GLM-5.3-Flash, so it drops into existing OpenAI-client code with a base URL and model-name change.
Use an API key from your DeepInfra dashboard.
Implementation Example
curl https://api.deepinfra.com/v1/openai/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $DEEPINFRA_API_KEY" \
-d '{
"model": "zai-org/GLM-5.3-Flash",
"messages": [
{
"role": "user",
"content": "Write a Python function to find the greatest common divisor (GCD) of two numbers."
}
]
}'| Parameter | Type | Description |
| model | String | Required. Set to zai-org/GLM-5.3-Flash. |
| messages | Array | Required. Supports text and multimodal (image) inputs. |
| response_format | Object | Optional. Use {“type”: “json_object”} for structured JSON output. |
| tools | Array | Optional. Define functions for agentic tool-calling tasks. |
| temperature | Float | Optional. Controls randomness (e.g., 0.2 for code, 0.8 for creative writing). |
See DeepInfra’s full API docs for streaming, vision inputs, and additional parameters.
Z.ai positions GLM-5.3-Flash as delivering high-tier intelligence at roughly one-tenth the cost of GLM-5.2. DeepInfra is currently running a 50% promotional discount on standard rates.
| Token Type | Standard Rate | Discounted Rate (Current) |
| Input | $0.15 | $0.075 |
| Output | $0.50 | $0.25 |
| Cached Input | $0.03 | $0.015 |
See the pricing guide for a full breakdown across scenarios and providers.
GLM-5.3-Flash pairs a 1M-token context window with an efficiency-focused MoE architecture, giving developers frontier-level coding, agentic, and multimodal reasoning capability at a fraction of typical cost. Whether you’re building autonomous software agents, analyzing visual documents, or processing large datasets, it’s available now on DeepInfra with an OpenAI-compatible API. See also: GLM-5.3-Flash Is Now on DeepInfra and the pricing and cost analysis.
Model Deprecation: Build LLM Apps That Last<p>Your model ID is the shortest-lived dependency in your stack and odds are it doesn’t have a maintenance schedule. On June 15, 2026, claude-sonnet-4-20250514 and claude-opus-4-20250514 stopped answering requests. Anthropic had posted the notice 62 days earlier. Teams with either string in a call site learned about it from an error rate, not an email. […]</p>
Best SaaS Tools and API Providers for MiMo-V2.5<p>As LLM architectures grow increasingly complex, the introduction of the MiMo-V2.5 series represents a significant step forward in multimodal capabilities and massive context handling. Integrating a model with a 1M-token context window and native multimodal support (image, video, audio, text) introduces substantial infrastructure considerations. For developers and enterprise architects, the priorities are clear: managing inference […]</p>
Best API for Kimi K2.5: Why DeepInfra Leads in Speed, TTFT, and Scalability<p>Kimi K2.5 is positioned as Moonshot AI’s “do-it-all” model for modern product workflows: native multimodality (text + vision/video), Instant vs. Thinking modes, and support for agentic / multi-agent (“swarm”) execution patterns. In real applications, though, model capability is only half the story. The provider’s inference stack determines the things your users actually feel: time-to-first-token (TTFT), […]</p>
© 2026 DeepInfra. All rights reserved.