DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3-Flash API Providers: Speed & Cost
Published on 2026.09.29 by DeepInfra
GLM-5.3-Flash API Providers: Speed & Cost

API Review Summary

MetricValue
Intelligence (Artificial Analysis Intelligence Index)42 — well above the open-weight median (18)
Speed55.9 output tokens/sec — slower than the median (85.7 t/s)
Latency (TTFT)3.14s — higher than the median (2.05s)
Cost (Z.ai first-party API)$0.15 / 1M input, $0.50 / 1M output; cache discount ~83%
Cost efficiency$0.25 cost per Intelligence Index task (weighted average)
Verbosity180M output tokens generated during the Intelligence Index eval — higher than median (140M)
Context window1,048,576 tokens (~1M; roughly 1,500 pages of text)
ModalitiesText + image input; text output
LicenseOpen weights, MIT (commercial use allowed)
Model size320B total parameters, 18B active (MoE)
Availability~20 API providers per Artificial Analysis (see provider benchmarks)

GLM-5.3-Flash — Best APIs

Each signal below is a reading from Artificial Analysis; the middle column says what it means for picking an API, and the right column is what to verify directly on deepinfra.com for this model.

Selection signal (Artificial Analysis)What it means for choosing an APIWhat to verify on DeepInfra
Slow output speed (55.9 t/s) vs. median (85.7 t/s)Prioritize providers that deliver higher real-world throughput for this model.Published throughput/benchmarking for this model, streaming behavior, region options.
Higher TTFT (3.14s) vs. median (2.05s)For interactive apps, provider-side routing and low-overhead serving matter as much as raw tokens/sec.Typical TTFT stats, streaming-first latency, any “fast start” or routing features.
Very verbose (180M tokens on the Index)Verbosity inflates output-token costs and end-to-end time — you want cost controls and stable rate limits.Support for max-output limits, response-length controls, and predictable rate limiting/quotas.
Low list pricing + large cache discount (~83%)Prompt caching and cache-hit passthrough can materially cut spend for repeated contexts, RAG, or system prompts.Whether DeepInfra supports prompt caching for this model, and how cache hits are billed/discounted.
$0.25 cost per Intelligence Index taskTotal cost depends on pricing + caching + provider add-ons — compare providers on the same workload.DeepInfra’s per-1M pricing (input/output + cache hit/write) vs. other providers.
1M-token context windowLong-context workloads need providers that reliably support the full window without truncation or excess latency.Max context supported in practice, per-request caps, long-context stability and timeouts.
Text + image inputIf you need vision, confirm the API accepts image inputs in the format you use.Image-input support, payload format (URL/base64/multipart), and any image size limits.
Open weights (MIT), MoE (320B total / 18B active)Open weights enable self-hosting, but a managed API can be simpler; MoE serving quality varies by provider implementation.Hardware transparency, reliability/SLA, and whether DeepInfra offers consistent deployments for this model.
~20 API providers availableUse provider benchmarks to pick the best combination of speed, TTFT, and price for your use case.Confirm GLM-5.3-Flash is listed and compare DeepInfra’s measured performance/pricing against the provider-benchmark list.

Released in August 2026 by Z.ai, GLM-5.3-Flash is a 320B-parameter Mixture-of-Experts (MoE) model with a 1M-token context window. Because the model is verbose and has slow baseline latency on the first-party API, the choice of hosting provider matters more than usual — DeepInfra is the overall recommended provider below for its full 1M context support, ~83% cache discount, and low $0.50/1M output pricing.

GLM-5.3-Flash made waves before anyone knew its name. For six days in August 2026 (August 20–26), a mystery model codenamed “Ox Alpha” topped usage charts on OpenRouter and OpenCode, offering near-unlimited free access while beating established benchmarks. Z.ai officially revealed GLM-5.3-Flash on August 26, 2026, by which point the model had already built a reputation in the wild.

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, combining sparse and linear attention — a first for an open-weight frontier model. Z.ai reports this hybrid design delivers roughly 3x lower attention compute and over 4x smaller KV cache versus the base GLM-5.3 model at long context lengths.

Z.ai says the model “starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency” — it isn’t a distilled or trimmed version of a flagship model, but a distinct model trained on a different corpus and optimized for coding and agentic workloads.

With MIT-licensed open weights on Hugging Face, GLM-5.3-Flash outperforms GLM-5.2 across benchmarks and real-world workloads at roughly one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. It supports text, image, and video input with a 1,048,576-token context window.

Provider Comparison Table

GLM-5.3-Flash is available across roughly 20 API providers, so picking the right host matters for mitigating its latency and maximizing cost-efficiency. The table below compares the top providers by speed, cost, and reliability — pulled from Artificial Analysis’s provider benchmarks plus each provider’s own published figures. Throughput and latency for these inference marketplaces shift often (day-to-day swings of 2–5x are common across aggregators), so treat the numbers below as directional and re-check live figures before publishing.

API ProviderBest ForOutput Speed (t/s)Blended Cost / 1M TokensNotable Feature
DeepInfraOverall recommendation––$0.03 cached input (~83% discount)
Z.ai (first-party)Native baseline55.9$0.103.14s TTFT
IncoRaw throughput~500–700–Claims 1.85x the next-fastest provider on AA’s leaderboard
BasetenEnd-to-end latency––3.23s p99 latency
Bitdeer AIBudget & batching––Live on Bitdeer AI Model Studio; competitively priced
Fireworks AIEnterprise reliability–$0.10Strong task-success track record on tool-calling workloads

Detailed Technical Analysis of API Providers

1. DeepInfra: Best Overall Provider

DeepInfra is a strong default for production deployments, balancing cost-efficiency, context handling, and reliability.

  • Input price: $0.15 per 1M tokens
  • Output price: $0.50 per 1M tokens
  • Cached input price: $0.03 per 1M tokens (~83% cache discount)
  • Context window support: full 1.0M tokens

Why it’s a good fit: GLM-5.3-Flash leans on its 1M context window for complex reasoning and agentic tasks, and DeepInfra fully supports that plus the cache discount. Because the model is verbose, DeepInfra’s flat $0.50/1M output price keeps long-horizon workflows predictable, and it offers a stable OpenAI-compatible endpoint for easy migration. Its real-world throughput varies across benchmarking sources and isn’t consistently the fastest of the providers here — the case for DeepInfra rests on price, context support, and API stability rather than raw speed.

2. Z.ai: Best for First-Party Native Features

As the model’s creator, Z.ai runs the baseline official API.

  • Output speed: 55.9 tokens/sec
  • TTFT: 3.14 seconds
  • Blended cost: ~$0.10 per 1M tokens (a 7:2:1 cache/input/output-weighted estimate)
  • Cost per task: $0.25 per Intelligence Index task

Use it when you need native, un-abstracted support for image inputs and raw reasoning parameters. Its 55.9 t/s speed is at the lower end for open-weight models this size (median: 85.7 t/s), so it’s not the best fit for latency-sensitive applications.

3. Inco: Best for Raw Throughput

Inco targets the model’s main bottleneck — its slow baseline generation speed. Inco has publicly stated throughput in the 500–700 tokens/sec range for GLM-5.3-Flash, which it describes as roughly 1.85x the next-fastest provider on Artificial Analysis’s leaderboard. For high-volume generation, large-scale extraction, or workflows where output speed drives user experience, it’s worth benchmarking directly against your workload — third-party throughput figures for this provider vary by source and testing window.

4. Baseten: Best for End-to-End Latency

Baseten optimizes routing and the “thinking” phase to minimize overall response time.

  • End-to-end p99 latency: 3.23 seconds (consistent across independent benchmarking sources)

End-to-end latency accounts for TTFT, reasoning time, and output speed together. For real-time agentic workflows — coding assistants, terminal use, the kind of task GLM-5.3-Flash is built for — a low, consistent p99 matters more than peak throughput, which is where Baseten positions itself.

5. Bitdeer AI: Best for Budget & Batch Processing

Bitdeer AI added GLM-5.3-Flash to its Model Studio and is positioned as a budget option for asynchronous batch processing — evaluation runs, quantitative analysis, or other workloads where turnaround time matters less than unit cost. Bitdeer hasn’t published a detailed public rate card for this model as of this writing, so confirm current pricing directly on their platform before committing to a workload.

6. Fireworks AI: Best for Task Success & Reliability

Fireworks AI is known for strict API contracts and high uptime, which matters for enterprise deployments.

  • Blended price: ~$0.10 per 1M tokens

For tool-calling and JSON-schema-heavy workflows — agentic pipelines where a dropped connection or malformed output can break a multi-step chain — Fireworks’ reliability track record is the draw, even where its raw speed or TTFT trails faster providers like Baseten or Inco.

Conclusion

GLM-5.3-Flash is a remarkably capable 320B-parameter model with frontier-level performance on coding and agentic benchmarks, but its verbosity and slow first-party baseline speed mean your choice of API provider matters more than usual.

For most developers and enterprises, DeepInfra is the top recommendation: full support for the 1M context window, an ~83% cache discount, and predictable $0.50/1M output pricing let you use the model’s reasoning capability without the cost unpredictability of less-optimized endpoints. See also the pricing and cost analysis and model documentation for this model.

Frequently Asked Questions

What is the best API provider for GLM-5.3-Flash? 

For most developers and enterprises, DeepInfra is the top recommendation. It supports the full 1M context window, offers an ~83% cache discount ($0.03 per 1M cached input tokens), and keeps output priced at $0.50 per 1M tokens.

What is the fastest API provider for GLM-5.3-Flash? 

Inco reports the highest raw throughput for this model, in the 500–700 tokens/sec range by its own published figures. For end-to-end latency specifically, Baseten’s 3.23s p99 is the strongest and most consistently corroborated figure across benchmarking sources.

How much does GLM-5.3-Flash cost to run? 

Cost varies by provider. The first-party Z.ai API blends to roughly 0.10per1Mtokens(0.25 per Intelligence Index task). Compare providers on your actual workload — blended-rate estimates can mislead if your traffic mix of cached/input/output tokens differs from the benchmark ratio.

Why is GLM-5.3-Flash called “Flash” if it’s not a distilled model? 

Unlike model families where “Flash” indicates a distilled or trimmed flagship, GLM-5.3-Flash is a separate model. Z.ai states it “starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency.” The name refers to efficiency gains from its hybrid sparse-linear attention architecture, not reduced capability.

What makes GLM-5.3-Flash different from GLM-5.3? 

GLM-5.3-Flash is the multimodal, high-efficiency variant priced at 0.15/0.50 per million tokens. The flagship GLM-5.3 is a text-only coding and cyber-focused model priced at roughly $1.40 input / $4.40 output per million — about ten times the price. For most non-coding workloads (chat, extraction, vision, summarization), GLM-5.3-Flash is the more practical choice.

Related articles
MiniMax-M2.5 API Benchmarks: Latency, Throughput & CostMiniMax-M2.5 API Benchmarks: Latency, Throughput & Cost<p>About MiniMax-M2.5 MiniMax-M2.5 is a state-of-the-art open-weights large language model released in February 2026. Built on a 230B-parameter Mixture of Experts (MoE) architecture with approximately 10 billion active parameters per forward pass, it features Lightning Attention and supports a context window of up to 205,000 tokens. The model uses extended chain-of-thought reasoning to work through [&hellip;]</p>
DeepSeek V4 Pro Is Now Available on DeepInfraDeepSeek V4 Pro Is Now Available on DeepInfra<p>DeepSeek released V4 Pro on April 24, 2026 — a 1.6 trillion-parameter Mixture of Experts model with 49 billion active parameters, a 1-million-token context window, and weights available on Hugging Face under an MIT license. On LiveCodeBench, the V4-Pro-Max reasoning variant scores 93.5 Pass@1, leading every model in the comparison set, including Gemini-3.1-Pro High at [&hellip;]</p>
Reliable JSON-Only Responses with DeepInfra LLMsReliable JSON-Only Responses with DeepInfra LLMs<p>When large language models are used inside real applications, their role changes fundamentally. Instead of chatting with users, they become infrastructure components: extracting information, transforming text, driving workflows, or powering APIs. In these scenarios, natural language is no longer the desired output. What applications need is structured data — and very often, that structure is [&hellip;]</p>