DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3 Is Now Available on DeepInfra
Published on 2026.10.03 by DeepInfra
GLM-5.3 Is Now Available on DeepInfra

GLM-5.3 shares its base model with GLM-5.2, so every performance gain comes from post-training. Z.ai scaled the reinforcement learning stack it already had and released the result on August 14, 2026. GLM-5.3 is a reasoning model built for complex software engineering and long-horizon agentic tasks. On Z.ai’s internal Code Bench, it improves 50% over GLM-5.2, and Z.ai reports it uses fewer output tokens to get there.

Cybersecurity is the bigger surprise. Multi-stage exploitation reasoning developed during post-training faster than Z.ai expected. GLM-5.3 tops the CyberGym benchmark and more than triples GLM-5.2 on ExploitGym. Working with security teams in China since GLM-5.2, Z.ai reports its models have found 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high severity issues. Findings are tracked on Z.ai’s public Security Disclosure Ledger.

The model pairs a 1M-token context window with three configurable reasoning effort levels. Some of its training environments mirror multi-day expert engineering work, so GLM-5.3 targets tasks that do not fit in a single prompt.

GLM-5.3 is now available on DeepInfra.

What Makes This Model Different

The base model is identical to GLM-5.2. Every gain comes from scaling the reinforcement learning stack Z.ai already had in place for that release. Three components carry it:

  • IndexShare handles long-context processing by reusing one indexer across sparse-attention layers, which cuts the compute cost of long contexts.
  • SAO is the reinforcement learning method for long-horizon tasks. It carries over from GLM-5.2 with compaction, which helps the gains hold on long runs as well as short ones.
  • slime is Z.ai’s open-source asynchronous framework for large-scale reinforcement learning.

None of the three is new in GLM-5.3. Z.ai added scale, with more environments, more varied tasks, longer episodes, and more compute, all aimed at the same base.

Z.ai also reports a more than 2.3x gain in end-to-end RL training throughput on long-horizon coding tasks. The figure is internal and cannot be verified from the outside. The larger result is a useful proof point that post-training scaling alone can drive a substantial capability jump on unchanged base weights.

Thinking is always on

GLM-5.3 always reasons, and Z.ai no longer supports disabling thinking. You control the compute budget through three reasoning_effort levels: low, high, and max. Max is the default and the level Z.ai recommends for coding. Teams that relied on a zero-thinking path in GLM-5.2 should move to low before migrating. GLM-5.3-Flash offers a lower per-token price for lighter workloads.

Coding performance

On Z.ai’s internal Code Bench, GLM-5.3 at max effort completes 34.5% of tasks against 23.4% for GLM-5.2, the roughly 50% gain the release leads with. Token cost looks better still. At high effort, GLM-5.3 scores 31.4% on about 50K output tokens, while Claude Opus 4.8 scores 29.5% on about 120K. GLM-5.3 still trails Claude Fable 5, which Z.ai puts at 39.5% at max effort.

Code Bench is a private benchmark, so these figures cannot be independently reproduced. The public benchmarks below are easier to check:

BenchmarkGLM-5.2GLM-5.3Kimi K3Opus 4.8GPT-5.6 SolFable 5 (w/ fallback)
Terminal Bench 2.181.088.288.385.088.888.0
Terminal Bench 3.04.628.317.421.134.633.7
DeepSWE v1.146.266.967.558.072.769.7
FrontierSWE67.578.1n/a66.5n/a88.2
PostTrainBench31.739.832.032.936.241.8

GPT-5.6 Sol leads Terminal Bench 2.1, Terminal Bench 3.0, and DeepSWE. Kimi K3 edges GLM-5.3 by 0.1 on Terminal Bench 2.1 and 0.6 on DeepSWE, then trails it by 10.9 points on Terminal Bench 3.0. Fable 5 leads FrontierSWE and PostTrainBench, with GLM-5.3 second on both.

The largest gain over GLM-5.2 is Terminal Bench 3.0, up from 4.6 to 28.3 on an unchanged base model. Z.ai publishes 16 benchmark rows in total. Two come from named outside evaluators: Artificial Analysis ran GDPval-AA v2, and Proximal ran FrontierSWE. Z.ai ran most of the other 14 rows itself, often through Claude Code 2.1.207.

Cybersecurity capability

Z.ai calls the cybersecurity gains emergent, meaning they developed faster than it anticipated during post-training. The company did add vulnerability-discovery data and environments on purpose. The surprise was the model planning complete exploitation chains across multiple stages, beyond spotting isolated flaws.

GLM-5.3 scores 84.5 on CyberGym, the best result among the models Z.ai lists. On ExploitGym it completes 105 tasks in the two-hour budget and 130 in the six-hour budget, up from 29 and 39. ExploitBench rises from 24.4 to 54.4, just over double.

The gap to the closed frontier remains wide. On the same ExploitGym budgets, GPT-5.6 Sol completes 216 and 293 tasks, and Fable 5 completes 181 and 247. Z.ai says as much: capability is growing fastest where it trails furthest.

Beyond benchmarks, Z.ai has run its models against real codebases with security teams in China since GLM-5.2. After expert review, screening, and deduplication, the work produced 2,436 findings across 269 projects, with 1,097 rated medium-to-high severity. Findings span system kernels, operating systems, browser engines, and network protocols, including Linux, WebKit, and FreeBSD. The oldest flaw dates back roughly 40 years. At launch, 53 findings were public with CVEs assigned and 2,383 remained under embargo, so the aggregate figures are Z.ai’s own until disclosure catches up. The ledger is public at cvd.z.ai.

Agentic performance

Z.ai reports open-weights state of the art for GLM-5.3 on Agents’ Last Exam CLI at 28.5, and it leads AutomationBench v1.0.6 at 48.2. It scores 1769 on GDPval-AA v2, the highest in Z.ai’s table, ahead of Kimi K3 (1682), Opus 4.8 (1588), and DeepSeek-V4 Pro (1590).

The model accepts text input only. It supports a 1,048,576-token context window, function calling, and JSON output. Together they suit long-horizon agent loops where context accumulation is the bottleneck.

Getting Started on DeepInfra

GLM-5.3 is available now on DeepInfra under the model identifier zai-org/GLM-5.3. A 38% promotional discount is currently active. Pricing is $0.563 per 1M input tokens, $2.50 per 1M output tokens, and $0.125 per 1M cached input tokens. List price without the promotion is $0.90, $4.00, and $0.20, so budget against list price for anything beyond the promotional window.

The Flex tier takes a further 20% off, bringing current rates to $0.45, $2.00, and $0.10. It applies the same reduction to input, output, and cache, and it stacks with the promotion. GLM-5.3 supports Flex and does not list a Priority tier. The model runs at fp4 quantization, supports the full 1,048,576-token context window, and has function calling and JSON output enabled out of the box.

A few API details are worth knowing before your first request. The reasoning_effort parameter accepts low, high, or max. It defaults to max when omitted or given any other value, so set it only when you want to trade quality for token efficiency. Pass clear_thinking=true in chat scenarios, since it defaults to false. To reproduce benchmark results, keep reasoning_effort at its default of max. The full parameter reference is on the GLM-5.3 API page.

DeepInfra offers an OpenAI-compatible API, so swapping your base URL and API key works with existing tooling. Pricing is usage-based with no minimums. The platform operates a zero-retention policy and holds SOC 2 and ISO 27001 certifications.

Here is a minimal working example:

cURL

curl "https://api.deepinfra.com/v1/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
  -d '{
    "model": "zai-org/GLM-5.3",
    "messages": [
      {
        "role": "user",
        "content": "Find the bug in this function and return a corrected version: def add(a, b): return a - b"
      }
    ],
    "reasoning_effort": "max"
  }'
copy

Python

from openai import OpenAI

client = OpenAI(
    api_key="$DEEPINFRA_TOKEN",
    base_url="https://api.deepinfra.com/v1/openai",
)

response = client.chat.completions.create(
    model="zai-org/GLM-5.3",
    messages=[{
        "role": "user",
        "content": "Find the bug in this function and return a corrected version: def add(a, b): return a - b",
    }],
    extra_body={"reasoning_effort": "max"},
)
print(response.choices[0].message.content)
copy

JavaScript

import OpenAI from "openai";

const openai = new OpenAI({
  apiKey: "$DEEPINFRA_TOKEN",
  baseURL: "https://api.deepinfra.com/v1/openai",
});

const response = await openai.chat.completions.create({
  model: "zai-org/GLM-5.3",
  messages: [{
    role: "user",
    content: "Find the bug in this function and return a corrected version: def add(a, b): return a - b",
  }],
  // the cast lets DeepInfra's "max" reasoning_effort value pass the SDK type check
  reasoning_effort: "max",
} as any);
console.log(response.choices[0].message.content);
copy

Run a request directly in the browser from the demo tab on the model page, or grab an API key and start integrating. If you need dedicated capacity, you can deploy GLM-5.3 as a private endpoint.

Conclusion

GLM-5.3 is a useful data point in the post-training scaling debate. Z.ai states its method plainly, same base weights and scaled reinforcement learning, and publishes the rows it loses alongside the rows it wins.

Token efficiency is the part to watch in production. If Z.ai’s Code Bench figures hold, edging Claude Opus 4.8 on about 50K output tokens against 120K would show up as a real cost difference.

Teams running long-horizon agents, automated security tooling, or complex software pipelines should run their own eval against whatever they use today. The Llama and Nemotron families on DeepInfra offer cheaper options for the easy steps in an agent loop. Qualifying startups can apply to DeepStart for up to 1 billion free tokens.

Related articles
Inference LoRA adapter modelInference LoRA adapter modelLearn how to inference LoRA adapter model.
Best DeepSeek-V4.1-Flash API Providers in 2026Best DeepSeek-V4.1-Flash API Providers in 2026<p>As LLM architectures grow increasingly sophisticated, deploying state-of-the-art models like DeepSeek-V4.1-Flash requires more than just a basic API wrapper. For engineering teams, the challenge lies in balancing time-to-first-token (TTFT), throughput, context caching, and enterprise-grade compliance. DeepSeek-V4.1-Flash offers remarkable capabilities—including a massive 1M+ token context window and Engram conditional memory—but unlocking its full potential depends heavily [&hellip;]</p>
Two-Tier AI Agents: Why You Need Two ModelsTwo-Tier AI Agents: Why You Need Two Models<p>Looking at the trace of a long-running coding or ops agent, you will see the same pattern. A significant share of the steps are mechanical: construct a tool call, check the output format, retry on a transient error, summarise a log, classify a result. None of that requires frontier reasoning. None of it requires a [&hellip;]</p>