DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It handles text input and output across a 1,048,576-token context window, and is one of several hundred models in DeepInfra’s catalog.
The model uses the same base as GLM-5.2. Every gain comes from post-training, which is where Z.ai concentrated its work on coding performance and on the balance between capability and token efficiency. If you need lower latency and lower cost on simpler work, GLM-5.3-Flash is the lighter sibling.
DeepInfra serves GLM-5.3 at fp4 with the full context window, under a zero-retention policy documented in the DeepInfra trust center. Other large open-weights models on the platform worth weighing against it include Qwen3.8-2.4T-A95B and DeepSeek-V4.1-Flash, which trades some reasoning depth for a much lower per-token price.
Model string: zai-org/GLM-5.3
Coding. Z.ai reports a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, and describes GLM-5.3 as the most capable open-weights model for coding. On public benchmarks, Z.ai claims open-source state of the art on Terminal Bench 3.0 and Agents’ Last Exam. Both claims are scoped to open-weights models. Closed frontier models still lead on several of these benchmarks, as the table below shows.
Cybersecurity. Z.ai describes cyber capability as developing faster than expected during post-training scaling. GLM-5.3 posts the highest CyberGym score among the models Z.ai compared, at 84.5. Its largest relative gains sit further up the exploitation chain: ExploitBench moves from 24.4 to 54.4, slightly more than double GLM-5.2.
Token efficiency. Z.ai positions GLM-5.3 as improving the balance between performance and token consumption relative to GLM-5.2. No specific reduction figure is published.
All figures below come from Z.ai’s published results for GLM-5.3. Z.ai ran these evaluations with its own harnesses and settings, so they are not interchangeable with numbers other labs publish for the same benchmarks. Most runs use Claude Code 2.1.207 at max reasoning effort. Z.ai’s full table also carries columns for DeepSeek-V4 Pro, Qwen3.8-Max and Fable 5, omitted here for width.
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | Opus 4.8 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 88.2 | 81.0 | 88.3 | 85.0 | 88.8 |
| Terminal Bench 3.0 | 28.3 | 4.6 | 17.4 | 21.1 | 34.6 |
| DeepSWE (v1.1) | 66.9 | 46.2 | 67.5 | 58.0 | 72.7 |
| NL2Repo | 58.0 | 48.9 | 58.0 | 69.7 | n/a |
| FrontierSWE | 78.1 | 67.5 | n/a | 66.5 | n/a |
| CyberGym | 84.5 | 77.2 | 80.0 | 78.1 | 83.6 |
| ExploitBench | 54.4 | 24.4 | 32.2 | 40.0 | 76.5 |
| Toolathlon Verified | 73.0 | 59.9 | 76.5 | 76.2 | 74.9 |
| AutomationBench (v1.0.6) | 48.2 | 26.2 | 46.7 | 41.0 | 45.8 |
| Agents’ Last Exam (ALE-CLI) | 28.5 | 23.8 | 27.6 | 25.7 | 28.6 |
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 57.9 | 64.5 |
| GDPval-AA v2 | 1769 | 1508 | 1682 | 1588 | 1730 |
GLM-5.3 posts the top score in this group on CyberGym, AutomationBench, FrontierSWE and GDPval-AA v2. GPT-5.6 Sol leads on Terminal Bench 2.1 and 3.0, DeepSWE, ExploitBench, Agents’ Last Exam and HLE with tools. Opus 4.8 leads on NL2Repo, and Kimi K3 on Toolathlon Verified and, by six tenths of a point, DeepSWE.
The gap against GPT-5.6 Sol is widest on ExploitBench, where GLM-5.3 scores 54.4 against 76.5. The gap is narrowest on Agents’ Last Exam, at 28.5 against 28.6.
Against GLM-5.2 the improvement is consistent and large on every benchmark listed, most dramatically on Terminal Bench 3.0, which moves from 4.6 to 28.3. To weigh GLM-5.3 against other open-weights options such as Qwen3-Max on price and capability, the model comparison tool puts them side by side.
Z.ai’s ExploitGym methodology rescales results by per-model tokens-per-second rates sourced from Artificial Analysis: 115 TPS for GLM-5.3, 40 TPS for Kimi K3 and 47 TPS for Qwen3.8-Max. These are inputs to a timeout calculation rather than measurements of DeepInfra’s serving speed. Throughput on DeepInfra depends on DeepInfra’s own fp4 deployment, so treat the 115 TPS figure as context for reading the benchmark, not as a service-level expectation.
A 25% promotional discount is currently active on GLM-5.3. The promotional column is what you pay today.
| Tier | Input | Output | Cached input |
|---|---|---|---|
| Standard, promotional | $0.90 | $3.00 | $0.15 |
| Standard, list price | $1.20 | $4.00 | $0.20 |
| Flex 0.8x, promotional | $0.72 | $2.40 | $0.12 |
All figures are per 1M tokens.
Two separate discounts are in play and they stack. The promotion takes 25% off list. The Flex service tier takes a further 20% off whatever the standard rate is at the time. Flex against list price alone would be $0.96 per 1M input tokens; the $0.72 figure reflects both discounts together. When the promotion ends, Flex rates rise with the standard rates.
Set service_tier to “flex” for the 20% discount, in exchange for slower, best-effort scheduling. A flex request can wait up to ten minutes for capacity before it runs or returns an HTTP 429, so use it for work you can retry: evaluations, data enrichment, asynchronous jobs. Leave service_tier unset for standard real-time scheduling. GLM-5.3 supports the Flex tier but not the Priority tier. Teams building something new can also apply for free inference credits through DeepStart.
The API is OpenAI-compatible. Point your existing OpenAI client at DeepInfra’s base URL and change the model name; the DeepInfra documentation covers the surrounding surface, including embeddings, speech and image endpoints.
Base URL: https://api.deepinfra.com/v1/openai
Endpoint: /chat/completions
Auth header: Authorization: Bearer $DEEPINFRA_TOKEN
import os
import requests
url = "https://api.deepinfra.com/v1/openai/chat/completions"
api_key = os.getenv("DEEPINFRA_TOKEN")
data = {
"model": "zai-org/GLM-5.3",
"messages": [
{"role": "user",
"content": "Explain the security implications of a buffer overflow."}
],
"clear_thinking": True,
"reasoning_effort": "max",
"max_tokens": 2048
}
response = requests.post(
url,
headers={"Authorization": f"Bearer {api_key}"},
json=data,
)
print(response.json())If you use the OpenAI SDK rather than raw HTTP, pass clear_thinking and reasoning_effort through extra_body, since they are not part of the standard OpenAI schema. The full API reference lists every accepted field.
| Parameter | Notes |
|---|---|
| model | Required. zai-org/GLM-5.3 |
| messages | Required. Roles: system, user, assistant. |
| reasoning_effort | Controls the thinking budget. Accepts low, high or max. Defaults to max when omitted or set to any other value. Keep max to reproduce published benchmark results. |
| clear_thinking | Defaults to false. Pass true for chat scenarios. |
| max_tokens | Maximum tokens to generate. See output limits below. |
| temperature | Sampling temperature, 0 to 2. Default 1.0. |
| top_p | Nucleus sampling threshold. Default 1.0. |
| stream | Set to true for server-sent events, terminating with [DONE]. |
| stop | Up to 4 sequences that halt generation. |
| n | Number of completion sequences. Default 1. |
| presence_penalty | Range -2.0 to 2.0. Default 0. |
| frequency_penalty | Range -2.0 to 2.0. Default 0. |
| response_format | Set to {“type”: “json_object”} for structured output. |
| tools, tool_choice | Function calling. |
| service_tier | “flex” for the discounted tier. Leave unset for standard. Priority is not supported on this model. |
| fail_fast | Return HTTP 429 immediately instead of queueing when the model is at capacity. Default false. |
Thinking budgets behave differently across model families on DeepInfra, so check the reasoning models guide before porting a prompt from another provider.
Maximum output tokens are model-dependent, with a hard cap of 16,384 tokens on most models. This is separate from the 1,048,576-token context window, which covers prompt and completion together.
For longer output, use response continuation: send a follow-up request with the previous response included as an assistant message, and the model picks up where it stopped. Continuation cannot push past the total context window, and a request that exceeds it returns a 400 error.
Several of Z.ai’s published evaluations run with output budgets well above DeepInfra’s default cap, in some cases 128K tokens. Reproducing those results through the shared serverless endpoint will require either response continuation or a private deployment.
GLM-5.3 stands as a powerhouse in the reasoning model landscape, offering unparalleled performance in coding and cybersecurity. By combining a massive 1M-token context window with sophisticated reasoning controls and industry-leading throughput, it enables developers to build the next generation of autonomous agents and security tools.
Next Steps:
Qwen3.8-27B Is Now Available on DeepInfra<p>The Qwen Team’s latest open-weight release, Qwen3.8-27B, is a 27-billion-parameter vision-language model. It goes deep on the tasks that trip up most models: multi-step agentic workflows, software engineering, and multimodal reasoning across images, documents, and video. On SWE-bench Pro, it scores 61.7, outpacing Opus4.6 Max’s 53.4. On QwenSWEBench it jumps from 49.3 on the previous […]</p>
Guaranteed JSON output on Open-Source LLMs.DeepInfra is proud to announce that we have released "JSON mode" across all of our text language models. It is available through the "response_format" object, which currently supports only {"type": "json_object"}
Our JSON mode will guarantee that all tokens returned in the output of a langua...
GLM-5.3 API Providers: Speed, Latency & Cost<p>GLM-5.3 API Review Summary GLM-5.3 is Z AI’s flagship coding and agentic reasoning model, released on August 14, 2026, and available on DeepInfra since launch. The model is built on the same ~753-billion parameter Mixture of Experts (MoE) base architecture as GLM-5.2, with all performance gains derived entirely from scaled post-training rather than architectural changes. […]</p>
© 2026 DeepInfra. All rights reserved.