DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3-Flash is a model from Z.ai in the GLM-5 family, released on August 26, 2026. It’s a natively multimodal model accepting text and image input (some listings, including OpenRouter, also note video), with text output. Both Artificial Analysis and DeepInfra describe it as a 320B-parameter mixture-of-experts model with 18B active parameters per inference; DeepInfra’s listing also notes fp4 precision. The model is available as open weights on Hugging Face under an MIT license, which changes the build-vs-buy conversation for teams that may want to self-host later.
What makes GLM-5.3-Flash stand out is that the economics aren’t coming at the expense of capability. Artificial Analysis gives it an Intelligence Index score of 42 (well above the 18 median for comparable open-weight models), while OpenRouter’s Artificial-Analysis-sourced figures put it at 41.8, alongside a Coding Index of 71.5 and an Agentic Index of 50.9. It also supports a large context window: 1M tokens per Artificial Analysis and DeepInfra, and 1,310,720 tokens on OpenRouter.
For developers and ML teams, the practical question isn’t whether the model is interesting, but where it makes operational sense. It has broad API availability: 20 providers per Artificial Analysis and 31 on OpenRouter, plus public and private deployment paths on DeepInfra. The tradeoff to understand upfront is that GLM-5.3-Flash is inexpensive but also notably slow and verbose: Artificial Analysis measured 55.9 output tokens/sec (versus an 85.4 median for comparable models) and 180M output tokens used during evaluation (versus a 140M median). If you’re cost-sensitive, building coding agents, or planning RAG and document-heavy systems where context length and cache pricing matter, this model deserves a serious pricing review rather than a quick skim.
GLM-5.3-Flash sits in an unusually competitive pricing band for a large open-weight reasoning model. Standard pricing is $0.15 input / $0.50 output per 1M tokens, DeepInfra lists a discounted rate of $0.075 / $0.25 / $0.015 cached, and OpenRouter shows provider prices ranging from about $0.045 / $0.14 / $0.01 up to $0.30 / $1.00 / $0.03. It’s best suited for developers building coding agents, long-context assistants, and document-heavy systems who want open weights, multimodal support, and the flexibility to choose between DeepInfra, OpenRouter-routed providers, and direct Z.ai access.
| Best For | Provider Recommendation | Why |
| Lowest price / cost-sensitive workloads | DeepInfra | DeepInfra lists one of the lowest explicitly stated discounted public rates in the research at $0.075 input, $0.25 output, and $0.015 cached input per 1M tokens. |
| RAG, document-heavy, or high-throughput use cases | DeepInfra | DeepInfra combines discounted pricing with a public API, a private deployment option, JSON output, function calling, multimodal input, and a roughly 1M-token context window. |
| Easiest onboarding / fastest time-to-first-call | DeepInfra | DeepInfra provides a public endpoint and an interactive demo interface, a clear low-friction starting point. |
| Proprietary or managed model access | Z.ai API | Artificial Analysis reports a 3.14s time to first token based on Z.ai’s own API, making direct vendor access the cleanest option for teams that want the model from its creator. |
| Broadest provider choice and routing flexibility | OpenRouter | OpenRouter lists 31 providers for GLM-5.3-Flash and offers Balanced, Nitro, Floor, and Exacto routing modes for different price, speed, and tool-calling priorities. |
| Best raw provider-side latency and throughput | OpenRouter provider marketplace | OpenRouter’s own figures show a best P50 latency of 0.43s and best-provider throughput of 117 tokens/sec, though results vary widely by provider. |
| Provider benchmarking before committing | OpenRouter | OpenRouter exposes provider-level pricing, uptime, availability, throughput, TTFT, and benchmark scores, useful for comparing DeepInfra, Baseten, CoreWeave, Crusoe, and others before locking in. |
Token pricing for GLM-5.3-Flash is simple on paper and easy to underestimate in practice. You pay separately for input tokens, output tokens, and sometimes cached input tokens. With this model, that split matters because the output side is much more expensive than the input side, and the model is also reported to be fairly verbose.
A token is a chunk of text, not a word. Short words may be one token; longer words, punctuation, code, JSON, and whitespace patterns can break differently. Images and other multimodal inputs are also converted into billable units behind the scenes, even if providers don’t always present them the same way. For budgeting, the only safe assumption is this: long prompts cost money, long answers cost more, and repeated context can be cheap or expensive depending on cache behavior.
| Token type | What it is | Why it matters |
| Input tokens | Your prompt, system message, chat history, tool schemas, retrieved documents, and any other text sent to the model | This is the baseline cost on every request. For GLM-5.3-Flash, input is relatively cheap, which helps for long-context and RAG workloads. |
| Output tokens | Everything the model generates back: answers, code, reasoning, JSON, tool arguments, summaries | This is where costs can jump. GLM-5.3-Flash charges more than 3x as much for output as input, and it’s reported to be verbose, so loose prompts can get expensive fast. |
| Cached input tokens | Reused prompt content that the provider recognizes and serves from cache instead of billing at full input price | This is the lever that makes repeated large prompts workable. If your app reuses system prompts, long instructions, or stable document prefixes, cache pricing can materially cut spend. |
For the headline rates commonly listed for GLM-5.3-Flash on Z.ai’s standard pricing:
The practical read: input is cheap enough to make large-context prompting realistic, output is the cost risk because it’s over 3x the input rate, and caching matters a lot for agent loops, long system prompts, and RAG pipelines that resend large shared context.
Artificial Analysis reports an 83% cache discount on Z.ai’s API and a blended rate of $0.10 per 1M tokens using a 7:2:1 mix of cache-hit/input/output tokens. That blended number is useful because it reflects how production systems often behave: repeated instructions, repeated conversation scaffolding, stable retrieved-context prefixes, and relatively smaller amounts of newly generated output.
The catch is that GLM-5.3-Flash is reported as very verbose: Artificial Analysis says it generated 180M output tokens during evaluation versus a 140M median for comparable models. If you don’t cap completions, tighten prompts, or control tool and result formatting, output spend can erase the benefit of cheap input.
For teams building with this model, token discipline matters more than usual:
The token math for GLM-5.3-Flash changes depending on where you run it. The model itself is the same, but provider pricing, cache treatment, and routing behavior can change your bill more than small benchmark differences ever will.
| Provider path | Token cost advantages | Token cost disadvantages |
| DeepInfra | One of the lowest clearly stated discounted public rates in the research: $0.075 input / $0.25 output / $0.015 cached input per 1M tokens. Strong fit for long-context apps that benefit from cheap cache reads. Public and private deployment options also help if you may later want more control over spend. | Discounted pricing may not be permanent. You still pay the same output-heavy ratio, so verbose generations aren’t magically cheap just because the input side is. |
| OpenRouter, discounted providers | Several OpenRouter-listed providers match or beat DeepInfra’s discounted rate; the cheapest listed at time of writing runs around $0.045 input / $0.14 output / $0.01 cache read. Useful if you want a single API while still getting reduced token pricing. | Routed platforms can make cost tracking messier. Actual provider selection, routing mode, and fallback behavior can affect what you pay and how predictable that spend is. |
| OpenRouter provider marketplace | Broad provider competition across 31 listed providers. Good if you want to shop for a price/performance sweet spot instead of taking one default endpoint. | Provider prices vary a lot: up to roughly $0.30 input / $1.00 output / $0.03 cache read on the high end (Cloudflare). If you don’t pin the provider or watch routing behavior, “same model” can quietly become a very different bill. |
| Direct Z.ai / standard list pricing | Straightforward reference pricing at $0.15 input / $0.50 output. Useful as the baseline when comparing whether routed or discounted options are actually saving you money. | Standard pricing is meaningfully higher than the discounted rates surfaced by DeepInfra and OpenRouter. If cost is your first filter, paying full list price is harder to justify unless you specifically want direct vendor access. |
| Self-hosting open weights | The model is MIT-licensed and weights are publicly available, so you aren’t locked into per-token API pricing forever. This can become attractive at sustained high volume or for private workloads with stable utilization. | Token prices disappear, but infrastructure costs show up instead. For a 320B MoE model with 18B active parameters, hosting economics aren’t casual: capacity planning, GPU utilization, latency tuning, and ops overhead replace the simplicity of per-token billing. |
The main provider-specific cost pattern is fairly clear:
If you care about token cost first, the rough order is: discounted DeepInfra or a discounted OpenRouter listing, then lower-cost OpenRouter providers, then standard list pricing, then self-hosting only after real utilization analysis.
If your workload is output-heavy, provider choice matters less than prompt control: saving a fraction of a cent on input doesn’t help much if the model keeps generating large answers at scale, and output tokens remain the most likely source of budget drift across every provider path.
If your workload is cache-friendly — repeated system prompts, long static instructions, reusable retrieved context, agent frameworks with stable scaffolding — provider choice matters more. That’s where $0.015 cached input versus full-price input starts to matter, especially under long-context usage where the prompt skeleton is much larger than the fresh user message.
DeepInfra stands out because it’s built for teams that care about price-performance, not just easy access. According to DeepInfra’s own account, the company was founded by a team with backend infrastructure experience at massive scale (they previously built and operated the infrastructure behind the imo messenger app), and it built its inference stack from the GPU hardware up rather than layering on top of a general-purpose cloud. For GLM-5.3-Flash, that infrastructure focus pairs with an aggressively discounted public rate. If you’re a developer, a high-volume API user, or a cost-conscious team trying to push serious token volume without overpaying, this is a provider worth shortlisting early.
| Model Name | Best Use Case | Context Window | Input Price (per 1M tokens) | Output Price (per 1M tokens) |
| GLM-5.3-Flash | Long-context assistants, coding agents, multimodal apps | 1,048,576 tokens | $0.075 | $0.25 |
Why this matters: on DeepInfra, GLM-5.3-Flash costs $0.075 per 1M input tokens and $0.25 per 1M output tokens, versus its standard listed rate of $0.15 input and $0.50 output — a straight 50% reduction on both sides — and cached input drops to $0.015 per 1M tokens. If your workload is large, repeated, or cache-friendly, that pricing delta adds up quickly.
For high-volume or cost-sensitive deployments, DeepInfra is the provider to test before you pay standard rates elsewhere. If GLM-5.3-Flash is on your shortlist, this is one of the clearest places to validate real-world economics fast, and you can confirm live availability and pricing on the model’s DeepInfra page before you commit production traffic.
Below are the kinds of workloads where DeepInfra makes GLM-5.3-Flash especially compelling: long-context, cache-friendly, agentic, and output-sensitive systems where a 50% discount on public API pricing is material. All scenarios use DeepInfra’s discounted rate ($0.075 input / $0.25 output / 0.015cached)andcompareagainststandardpricing(0.15 input / $0.50 output / $0.03 cached).
A developer team is building an internal support assistant that sends a long system prompt, retrieval instructions, and stable policy text on most requests. This is exactly the kind of workload where DeepInfra’s low cached-input pricing helps, because the reusable prompt scaffold doesn’t need to be billed at full input rates every time.
| Metric | Value |
| Volume | 100M cached input + 20M fresh input + 10M output tokens / month |
| Model | GLM-5.3-Flash |
| Provider | DeepInfra |
| Monthly Cost | $5.50 |
Cost breakdown: cached input 100M × $0.015 = $1.50, fresh input 20M × $0.075 = $1.50, output 10M × $0.25 = $2.50. Total: $5.50.
Comparison: the same token mix at standard pricing ($0.15 input, $0.50 output, $0.03 cached) would cost $11.00, so DeepInfra cuts that to $5.50.
This is one of the clearest fits for GLM-5.3-Flash. Coding and agentic frameworks are a large share of real-world usage for models in this class, and DeepInfra gives you the low-cost path if you want that capability without paying standard rates. The model’s long context also helps when the agent needs to inspect multiple files, tool schemas, and previous steps in one pass.
| Metric | Value |
| Volume | 50M input + 30M output tokens / month |
| Model | GLM-5.3-Flash |
| Provider | DeepInfra |
| Monthly Cost | $11.25 |
Cost breakdown: input 50M × $0.075 = $3.75, output 30M × $0.25 = $7.50. Total: $11.25.
Why DeepInfra looks strong here: discounted rates, a public endpoint for fast testing, a private deployment option if the workflow later needs tighter control, and function calling and JSON outputs for structured agent loops.
Comparison: the same workload at standard pricing would cost $22.50, so DeepInfra saves $11.25 per month.
A team processes large contracts, technical manuals, or audit files and asks for compact structured summaries. This is where DeepInfra’s combination of a roughly 1M-token context window, multimodal support, function calling, and discounted input pricing is unusually attractive. Input-heavy applications benefit more than usual because GLM-5.3-Flash input is already inexpensive, and DeepInfra halves it again.
| Metric | Value |
| Volume | 200M input + 20M output tokens / month |
| Model | GLM-5.3-Flash |
| Provider | DeepInfra |
| Monthly Cost | $20.00 |
Cost breakdown: input 200M × $0.075 = $15.00, output 20M × $0.25 = $5.00. Total: $20.00.
Why DeepInfra is a good match: cheap input for document-heavy workloads, low enough output pricing to make summary generation practical, no need to split aggressively around small context windows, and a public API now with a private endpoint later if the pipeline becomes business-critical.
Comparison: at standard pricing, this workload would cost $40.00, so DeepInfra reduces it to $20.00.
Suppose you’re building a developer-facing app that accepts screenshots, diagrams, or mixed text-and-image inputs and returns JSON for downstream systems. DeepInfra is a strong fit because it exposes multimodal input, JSON response formatting, and function calling on the same discounted endpoint.
| Metric | Value |
| Volume | 80M input + 15M output tokens / month |
| Model | GLM-5.3-Flash |
| Provider | DeepInfra |
| Monthly Cost | $9.75 |
Cost breakdown: input 80M × $0.075 = $6.00, output 15M × $0.25 = $3.75. Total: $9.75.
Why DeepInfra stands out: native multimodal support at a discounted rate, structured output support that helps contain the model’s verbosity, and practical economics for extraction pipelines where repeated calls add up quickly.
Comparison: at standard pricing, the same workload would cost $19.50, versus $9.75 on DeepInfra.
This scenario fits teams running many repeated agent calls with stable system prompts, tool definitions, and orchestration scaffolding. DeepInfra is particularly strong here because cached input falls to $0.015 per 1M tokens, cheap enough to make large repeated prompt frames economically viable.
| Metric | Value |
| Volume | 500M cached input + 100M fresh input + 50M output tokens / month |
| Model | GLM-5.3-Flash |
| Provider | DeepInfra |
| Monthly Cost | $27.50 |
Cost breakdown: cached input 500M × $0.015 = $7.50, fresh input 100M × $0.075 = $7.50, output 50M × $0.25 = $12.50. Total: $27.50.
Why this is a DeepInfra-strength use case: repeated orchestration context gets billed at the lowest stated cached rate, a public endpoint is enough to validate the economics quickly, and a private deployment path gives a clean upgrade route if traffic or compliance requirements grow.
Comparison: at standard pricing, the same workload would cost $55.00, so DeepInfra saves $27.50.
Across these scenarios, the pattern is consistent: DeepInfra is strongest when you want GLM-5.3-Flash for long-context assistants, coding agents, multimodal pipelines, and cache-heavy production systems without paying standard list pricing. The biggest caveat is still output control. Because GLM-5.3-Flash is described as verbose, the best savings come when you pair DeepInfra’s lower rates with tight prompting, structured outputs, and conservative completion limits.
Choosing a provider for GLM-5.3-Flash is less about picking a winner and more about matching your workload’s actual cost structure to the right pricing model. The model’s combination of long context, multimodal input, function calling, and MIT-licensed weights gives you real flexibility, but that flexibility only translates into savings if you understand where your token spend actually concentrates: input, output, or cached context.
The two criteria that matter most in practice are output verbosity and cache hit rate. GLM-5.3-Flash generates more output than comparable models, so the output price you negotiate up front has an outsized effect on your monthly bill. If your workload is cache-friendly — stable system prompts, shared orchestration scaffolding, repeated document prefixes — cached input pricing at $0.015 per 1M tokens on DeepInfra is the lever that can cut your effective cost well below any blended estimate. If your workload is output-heavy with low cache reuse, prompt discipline matters more than provider selection.
For most developers evaluating this model seriously, DeepInfra is a solid place to start: the discounted public rates are among the lowest clearly stated in the research, the endpoint supports multimodal input, JSON outputs, and function calling, and the interactive demo lets you test the model before writing a line of integration code. The API documentation covers what you need for a production call, and if GLM-5.3-Flash isn’t the right fit, DeepInfra’s full model catalog gives you a direct comparison against other options at the same infrastructure layer. Run your actual token mix against these numbers, check your cache-hit assumptions, and you should know within an hour whether this model and provider combination makes sense for your build.
GLM-4.7-Flash API Benchmarks: Latency, Throughput & Cost<p>About GLM-4.7-Flash GLM-4.7-Flash is Z.AI’s open-weights reasoning model released in January 2026. Built on a Mixture-of-Experts (MoE) Transformer architecture, it features 30 billion total parameters with only ~3 billion active per inference — making it exceptionally efficient for its capability class. The model is designed as a lightweight, cost-effective alternative to Z.AI’s flagship GLM-4.7, optimized […]</p>
GLM 5.2 vs Claude Opus 4.8: Pricing the Task, Not the Token<p>Every GLM 5.2 vs Claude Opus 4.8 comparison lands in the same place. Opus wins most coding benchmarks, GLM costs a fraction as much, pick according to your budget. That framing takes the price cards at face value, but it’s misleading. Price a finished unit of work instead of a million tokens and the gap […]</p>
How Open Source AI Is Closing the Gap<p>At the end of 2023, the gap between open-weight and closed-source AI models was real and easy to describe. If you wanted the best performance on reasoning, language understanding, or multi-step problem solving, you paid for a proprietary API. Open models were useful, capable for many tasks, and dramatically cheaper to run but they were […]</p>
© 2026 DeepInfra. All rights reserved.