DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3-Flash API Is Now on DeepInfra
Published on 2026.09.28 by DeepInfra
GLM-5.3-Flash API Is Now on DeepInfra

Z.ai’s GLM-5.3-Flash activates just 18 billion of its 320 billion parameters at inference time, and on Z.ai’s reported results it outscores Claude Opus 4.8 on DeepSWE v1.1 (63.4 vs. 58.0), a demanding software engineering benchmark. It’s the first natively multimodal model in the GLM-5 series, combining text, image, and video input with a one-million-token context window, at roughly one-tenth the price of GLM-5.2.

That efficiency comes from a purpose-built hybrid architecture that pairs sparse and linear attention, cutting attention compute by 3× and shrinking the KV cache by 4.4× compared with the full GLM-5.3, a meaningful difference when you’re serving long contexts at scale.

Before its official release, Z.ai tested the model anonymously as “ox-alpha” on OpenCode and OpenRouter, where, according to Z.ai, it became the most popular model of the week. It was trained on a 30-trillion-token multimodal corpus. GLM-5.3-Flash is now available in DeepInfra’s model catalog.

What Makes This Model Different

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, and it starts from a newly trained base model rather than an update to an earlier checkpoint. The architecture introduces two notable mechanisms: Manifold-Constrained Hyper-Connections (mHC), which improve scaling efficiency, and IndexPool, which compresses four indexer key vectors into one through weighted pooling to reduce latency and memory overhead at long contexts. The headline change is the shift to a hybrid sparse and linear attention design, a first for the GLM series, which makes million-token context lengths operationally tractable rather than just technically possible.

That hybrid attention pays off measurably at scale. Compared with the full GLM-5.3, GLM-5.3-Flash uses 3.0× less attention compute and a 4.4× smaller KV cache. Compared with GLM-4.5, it nearly halves both activated parameters (18B vs. 32B) and layers (45 vs. 92), while total parameter count stays similar (320B vs. 355B). The result is a 320B mixture-of-experts model that only runs 18B parameters per forward pass.

On benchmarks reported by Z.ai, the gains over GLM-5.2 are substantial across the board:

BenchmarkGLM-5.3-FlashGLM-5.2Claude Opus 4.8
DeepSWE v1.163.446.258.0
AutomationBench v1.0.648.826.241.0
Toolathlon Verified78.459.976.2
NL2Repo56.348.969.7
GDPval-AA v2177315041582

On Z.ai’s internal Code Bench v1.0 at max effort, it scores 29.0 against Claude Opus 4.8’s 29.5. All figures are self-reported by Z.ai and the test harnesses differ by benchmark, so see the footnotes on the model card before comparing across rows.

The native vision capability is new territory for the GLM series. GLM-5.3-Flash accepts image and video inputs out of the box, and Z.ai reports 80.5 on MMVU, 89.4 on CharXiv Reasoning with tools, 78.0 on Chartography with tools, 77.8 on MVBench, and 62.4 on OfficeQA Pro. That makes it a reasonable candidate for document and chart analysis, GUI agents, and computer use workflows, not just text-in/text-out tasks. You can explore multimodal models on DeepInfra to see how it compares to others in that category.

Getting Started on DeepInfra

GLM-5.3-Flash is available under the model identifier zai-org/GLM-5.3-Flash. Standard pricing is $0.15 per 1M input tokens and $0.50 per 1M output tokens, with cached tokens at $0.03 per 1M. A 50% discount is currently active, bringing those figures down to $0.075, $0.25, and $0.015 respectively; the pricing page always shows current rates. The full 1,048,576-token context window is supported, along with JSON mode, function calling, and multimodal inputs. If you need dedicated capacity, you can deploy a private endpoint.

DeepInfra exposes the model through an OpenAI-compatible API: no infrastructure to provision, no model weights to manage, and usage-based billing from the first token. If you’re already using the OpenAI SDK, the only things that change are the base URL and the model name. The API page for this model lists the available parameters, including the reasoning_effort field (low, high, or max) for trading latency against accuracy, and the DeepInfra docs cover the rest of the API.

DeepInfra operates a zero-retention policy on inference data and is both SOC 2 and ISO 27001 certified.

Here’s everything you need to make your first call.

cURL:

curl "https://api.deepinfra.com/v1/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
  -d '{
    "model": "zai-org/GLM-5.3-Flash",
    "messages": [
      {
        "role": "user",
        "content": "Explain the tradeoffs between sparse and linear attention in one paragraph."
      }
    ]
  }'
Python:
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DEEPINFRA_TOKEN"],
    base_url="https://api.deepinfra.com/v1/openai",
)

response = client.chat.completions.create(
    model="zai-org/GLM-5.3-Flash",
    messages=[
        {
            "role": "user",
            "content": "Explain the tradeoffs between sparse and linear attention in one paragraph.",
        }
    ],
)

print(response.choices[0].message.content)
JavaScript:
import OpenAI from "openai";

const openai = new OpenAI({
  apiKey: process.env.DEEPINFRA_TOKEN,
  baseURL: "https://api.deepinfra.com/v1/openai",
});

const response = await openai.chat.completions.create({
  model: "zai-org/GLM-5.3-Flash",
  messages: [
    {
      role: "user",
      content:
        "Explain the tradeoffs between sparse and linear attention in one paragraph.",
    },
  ],
});

console.log(response.choices[0].message.content);
copy

Grab your API key from the DeepInfra dashboard, then head to the model demo page to run a quick test or review the full spec before you integrate.

Wrapping Up

GLM-5.3-Flash is a technically interesting result: a 320B mixture-of-experts model that runs 18B parameters per forward pass, handles vision natively, and, on Z.ai’s reported numbers, holds its own against frontier models on software engineering benchmarks at a fraction of the cost. The architectural choices here aren’t cosmetic. The hybrid attention design and IndexPool compression are what make million-token multimodal inference practical to serve.

That combination opens up a real category of agents (code assistants with full repo context, GUI automation pipelines, document-heavy workflows) that were often too expensive to build reliably before. If you want to see how it handles your workload before committing, the live demo is a reasonable starting point, and teams that need guaranteed capacity can set up a private endpoint. To budget an agent workload, read our take on why cost per task matters more than cost per token and how two-tier agents pair a frontier model with a cheaper one. You can also browse the rest of our text generation models.

Related articles
Design Your Next Website With AI: A Prompting Guide for Ming-ImageDesign Your Next Website With AI: A Prompting Guide for Ming-ImageLearn how to prompt Ming-Image, an open-weight model tuned for UI design, to turn a one-line brief into a real website mockup.
Seed Anchoring and Parameter Tweaking with SDXL Turbo: Create Stunning Cubist ArtSeed Anchoring and Parameter Tweaking with SDXL Turbo: Create Stunning Cubist ArtIn this blog post, we're going to explore how to create stunning cubist art using SDXL Turbo using some advanced image generation techniques.
Model Deprecation: Build LLM Apps That LastModel Deprecation: Build LLM Apps That Last<p>Your model ID is the shortest-lived dependency in your stack and odds are it doesn’t have a maintenance schedule. On June 15, 2026, claude-sonnet-4-20250514 and claude-opus-4-20250514 stopped answering requests. Anthropic had posted the notice 62 days earlier. Teams with either string in a call site learned about it from an error rate, not an email. [&hellip;]</p>