DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3 is a large-scale reasoning model from Z.ai, released on August 18, 2026. Artificial Analysis tracks it as GLM-5.3 (max), emphasizing the reasoning variant, while OpenRouter lists it as Z.ai: GLM 5.3 and DeepInfra hosts it as zai-org/GLM-5.3. The model is open-weights, uses a Mixture of Experts architecture with 753 billion total parameters and 40 billion active parameters per token, and offers roughly a 1 million token context window, depending on provider presentation. It is text-in, text-out, supports structured outputs and function calling in hosted deployments, and its weights are publicly available under the GLM-5.3 License, which allows commercial use with restrictions.
What makes GLM-5.3 interesting is that it combines strong benchmark positioning with the operational tradeoffs teams actually feel in production. Artificial Analysis gives it an Intelligence Index score of 45, well above the median for comparable open-weight models in its size class, but also calls it particularly expensive, slower than average, and very verbose. DeepInfra’s documentation adds the more practical angle: GLM-5.3 uses the same base model as GLM-5.2, with improvements coming from post-training, and Z.ai’s own evaluations show meaningful gains on coding, agent, and cyber tasks versus GLM-5.2, while also comparing it against models like Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, Opus 4.8, Fable 5, and GPT-5.6 Sol. In other words, this is not a cheap general-purpose default. It is a model you reach for when long-horizon software work, tool use, or very large context windows justify higher output spend.
For technical decision-makers, the real question is not whether GLM-5.3 is good. The research already makes clear that it is competitive. The real question is whether its price profile, verbosity, and provider-specific performance line up with your deployment constraints.
GLM-5.3 is best viewed as a high-end open-weights reasoning model with a wide pricing spread depending on where you buy it. Research shows list pricing around $1.40/M input and $4.40/M output, while discounted hosted options like DeepInfra come in at $0.90/M input and $3.00/M output, and OpenRouter reports access across 30+ providers. It is best suited for teams that need long-context coding, tool use, or agent workflows and want a serious alternative to models like GLM-5.2, Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, Opus 4.8, Fable 5, and GPT-5.6 Sol without giving up open-weight deployment options.
| Best For | Provider Recommendation | Why |
| Lowest price / cost-sensitive workloads | DeepInfra Flex Tier | At $0.72/M input, $2.40/M output, and $0.12/M cached input, it is the lowest clearly documented price in the research. |
| Proprietary or managed model access | DeepInfra Public or Private Endpoint | DeepInfra offers a public endpoint, private endpoint deployment, zero data retention, and support for JSON plus function calling. |
| Easiest onboarding / fastest time-to-first-call | OpenRouter | OpenRouter provides OpenAI-compatible access, automatic failover, and routing across 30 providers, which reduces integration friction. |
| RAG, document-heavy, or high-throughput use cases | DeepInfra Standard Endpoint | DeepInfra combines a 1,048,576-token context window with discounted cached input pricing at $0.15/M, which is practical for large repeated-context workloads. |
| Fastest observed latency | OpenRouter provider at 0.37s TTFT / 120 tps | OpenRouter’s provider-level data shows the best observed P50 latency and throughput for GLM-5.3 in the research. |
| Open-weights deployment flexibility | Self-hosted weights from zai-org/GLM-5.3 | GLM-5.3’s weights are publicly available, so teams that want infrastructure control are not locked into a single hosted vendor. |
Token pricing is where GLM-5.3 stops being an abstract “strong reasoning model” and starts affecting your budget.
A token is a small chunk of text. In practice:
Here’s the short version of what you are usually paying for:
| Token type | What it is | Why it matters |
| Input tokens | Everything you send to the model: system prompt, user prompt, tool definitions, chat history, retrieved docs | Large contexts get expensive fast, especially with a 1M+ token window that invites abuse |
| Output tokens | Everything the model generates back | Usually the biggest cost driver for GLM-5.3 because it tends to produce long answers and long reasoning-heavy completions |
| Cached input tokens | Reused input context that a provider discounts after the first use | This is where RAG, agent memory, and repeated large prefixes can become much cheaper |
| Max completion tokens | The cap on how many output tokens the model is allowed to generate | A practical guardrail against runaway responses and surprise invoices |
| Reasoning tokens | Extra internal or reasoning-heavy generation associated with extended thinking behavior | You do not always see these broken out separately in billing, but you feel them in output length and latency |
A few practical implications for GLM-5.3:
The base list price for GLM-5.3 is fairly consistent across sources: about $1.40/M input and $4.40/M output. The useful differences show up in discounts, cached pricing, routing options, and whether you can access lower-cost variants without changing your integration.
| Provider / route | Token cost advantage | Token cost disadvantage | Best fit |
| DeepInfra Standard Endpoint | Lower than list price at $0.90/M input, $3.00/M output, $0.15/M cached input | Still not cheap if your app produces long completions; output cost remains the main driver | Teams that want predictable managed pricing and expect repeated long-context prompts |
| DeepInfra Flex Tier | Lowest documented price in the research at $0.72/M input, $2.40/M output, $0.12/M cached input | Lower price usually comes with tradeoffs in scheduling flexibility or capacity behavior typical of flex-style tiers | Cost-sensitive batchy workloads, offline jobs, and agent tasks that can tolerate some variability |
| OpenRouter discounted routing | Blended discounted pricing around $0.8775/M input, $2.97/M output, $0.1755/M cached; easy access to many providers without rewriting code | Cached pricing is worse than DeepInfra’s documented cached rate; provider prices and performance vary a lot underneath the same model name | Teams optimizing for convenience, routing, and failover more than absolute lowest token cost |
| OpenRouter fastest observed provider | Can pair the same model with very strong latency, including reported best cases around 0.37s TTFT and 120 tps | Fastest observed provider pricing can be as high as $1.40/M input and $4.40/M output, so speed may cost materially more per generated token | User-facing apps where responsiveness matters more than squeezing every cent |
| OpenRouter higher-priced providers | Extra provider choice, regional options, and fallback paths | Input pricing reaches $2.10/M on some providers, which is a steep premium for the same underlying model | Only sensible when you need a specific provider characteristic beyond price |
| Z.ai / list-price baseline as tracked by Artificial Analysis | Useful reference point for true undiscounted economics; includes an estimated 81% cache discount and a $0.90/M blended rate under a 7:2:1 cache/input/output mix | By raw list rates, GLM-5.3 is classified as expensive versus similar open-weight models; $4.40/M output is the part that hurts | Budget modeling and apples-to-apples benchmarking against other models |
| Self-hosted weights | You can avoid per-token API markups and design your own caching, batching, and context management strategy | Token cost turns into infrastructure cost, ops time, and utilization risk; “cheaper” disappears quickly if throughput is poor | Teams with stable high volume, infra experience, and a real reason to own the stack |
Practical takeaways by workload type:
If you are estimating GLM-5.3 cost, use this mental model:
That order saves more money than arguing about prompt wording for three days.
DeepInfra is the kind of provider that makes sense when you care about infrastructure economics, not just model access. It runs on bare-metal infrastructure, which matters because cutting out extra virtualization layers can help improve efficiency and keep pricing lower while still supporting serious throughput. In practice, DeepInfra’s machine learning infrastructure is typically 50–80% cheaper than major cloud competitors, which is exactly why it stands out for developers, high-volume API users, and cost-conscious teams that expect real production traffic. For GLM-5.3 specifically, it combines aggressive token pricing with practical deployment options like public access, private endpoints, zero retention, and a Flex tier for even lower-cost runs. You can also browse the full catalog of text generation models to compare GLM-5.3 against similar open-weight options.
| Model Name | Best Use Case | Context Window | Input Price (per 1M tokens) | Output Price (per 1M tokens) |
| GLM-5.3 Standard | Managed production workloads with strong reasoning and long-context support | 1,048,576 tokens | $0.90 | $3.00 |
| GLM-5.3 Flex Tier | Lowest-cost batch, offline, or flexible-scheduling workloads | 1,048,576 tokens | $0.72 | $2.40 |
| GLM-5.3 Public Endpoint | Fastest path to API access on DeepInfra | 1,048,576 tokens | $0.90 | $3.00 |
| GLM-5.3 Private Endpoint | Teams needing dedicated deployment and tighter operational control | 1,048,576 tokens | $0.90 | $3.00 |
| GLM-5.3 JSON / Function Calling | Structured-output apps and tool-using agent systems | 1,048,576 tokens | $0.90 | $3.00 |
Why this matters: On DeepInfra, GLM-5.3 is priced at $0.90/M input and $3.00/M output, with a Flex option at $0.72/M input and $2.40/M output. That is materially below the model’s standard listed pricing of $1.40/M input and $4.40/M output, so if you expect sustained volume, the savings show up quickly. For cost-sensitive reasoning workloads, DeepInfra is one of the clearest ways to bring GLM-5.3 spend down without changing models.
If GLM-5.3 is on your shortlist and token volume is going to be meaningful, DeepInfra is the provider to price first. It is especially compelling when you want managed access without paying full list rates.
These examples are where DeepInfra looks especially strong for GLM-5.3: long-context workloads, repeated-context systems, and agent-heavy jobs where you want managed access without paying full list pricing.
A team builds an internal coding copilot that sends large repository context, architectural notes, and tool schemas into GLM-5.3 for code review, refactors, and implementation planning. This is a classic fit for DeepInfra because GLM-5.3 is positioned for complex software engineering, and DeepInfra undercuts standard list pricing while keeping the full 1,048,576-token context window.
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 10M input / 5M output per month | GLM-5.3 | DeepInfra Standard | 10M | 5M | $24.00 |
Cost breakdown:
Same workload on a more expensive provider: at standard list pricing of $1.40/M input and $4.40/M output, the same usage would cost $36.00/month, or $12.00 more.
A developer team runs a support or internal search assistant over large docs, runbooks, API references, and architecture guides. The system repeatedly reuses big prompt prefixes and retrieved context. This is one of DeepInfra’s best cases because cached input is explicitly priced at $0.15/M on Standard and $0.12/M on Flex.
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 20M fresh input / 100M cached input / 10M output per month | GLM-5.3 | DeepInfra Standard | 20M input + 100M cached | 10M | $49.00 |
Cost breakdown:
Same workload on a more expensive provider: at $1.40/M input, $0.26/M cached, and $4.40/M output, this would cost $98.00/month, or $35.00 more.
This is the kind of workload where the Flex tier is hard to ignore. If you are processing backlog tasks overnight, migrating services, generating test fixes, or running multi-step software agents in batch, Flex gives you the lowest documented GLM-5.3 pricing in the research.
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 100M input / 40M output per month | GLM-5.3 | DeepInfra Flex | 100M | 40M | $168.00 |
Cost breakdown:
Same workload on a more expensive provider: at standard list pricing of $1.40/M input and $4.40/M output, the same batch job volume would cost $316.00/month, or $148.00 more.
A team is building a backend service that needs structured outputs, function calling, and zero-retention handling for internal automation. DeepInfra is attractive here because those capabilities are documented directly for GLM-5.3, so you are not paying extra list pricing just to get the hosted features developers usually need. Teams building their own integrations often reference DeepStart’s production-ready model examples when wiring up structured output flows.
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 30M input / 15M output per month | GLM-5.3 | DeepInfra Standard | 30M | 15M | $72.00 |
Cost breakdown:
Same workload on a more expensive provider: at $1.40/M input and $4.40/M output, the same workload would cost $108.00/month, or $36.00 more.
Some teams want GLM-5.3 behind a private endpoint for tighter operational control without giving up managed inference. DeepInfra is notable here because private endpoint deployment is explicitly supported, but pricing remains aligned with its documented GLM-5.3 endpoint rates in the research summary rather than jumping to the model’s higher list baseline.
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 50M input / 25M output per month | GLM-5.3 | DeepInfra Private Endpoint | 50M | 25M | $120.00 |
Cost breakdown:
Same workload on a more expensive provider: at $1.40/M input and $4.40/M output, this would cost $180.00/month, or $60.00 more.
For GLM-5.3, DeepInfra’s strengths are pretty clear:
If you already know GLM-5.3 is the right model, DeepInfra is the provider that most directly improves the economics for developer-heavy production use. If your workflows also touch adjacent capabilities, teams sometimes pair it with a companion route like GLM-5.3-Flash for voice-oriented tasks that do not require the full reasoning depth of the larger model.
Choosing a provider for GLM-5.3 is not a minor implementation detail. The spread between list pricing and what you actually pay on a discounted managed endpoint is wide enough to change whether a production workload is economically viable, and the model’s verbosity means output token costs accumulate faster than they do on more terse alternatives. Provider choice shapes your latency floor, your caching economics, and whether features like function calling and zero retention are available without extra configuration.
The clearest decision criteria are output pricing, cached input rates, and deployment flexibility. GLM-5.3 generates long completions by design, so a lower output rate compounds across every request. If your workload involves repeated large context — RAG pipelines, agent memory, stable system prompts — cached input pricing deserves the same attention as the headline input rate. And if you need private endpoints, JSON output, or zero retention, those need to be confirmed at the provider level before you commit to an integration. Developers who want to inspect the exact request format can review the GLM-5.3-Flash API documentation for a close analogue of the Standard endpoint’s interface, and can browse the broader DeepInfra model catalog to compare adjacent options.
For teams running batch jobs, offline agents, or any workload that can tolerate flexible scheduling, DeepInfra’s Flex tier offers the lowest documented pricing in this guide without requiring a model switch. That is a real option worth pricing before defaulting to standard endpoint rates. If you are also evaluating a lighter alternative for tasks that do not need full reasoning depth, GLM-5.3-Flash is worth a look alongside the main model, and for teams still weighing generational tradeoffs, the earlier GLM-4.7 MoE model remains available as a cheaper baseline for agentic coding and tool use.
The research is clear that GLM-5.3 is a capable model for the workloads it targets. What is less obvious until you run the numbers is how much provider selection moves the final invoice. Start with DeepInfra, run your actual prompt mix against the Flex and Standard tiers, and let real token counts drive the decision.
Kimi K2.6 Model Overview: Architecture, Features & Capabilities<p>Kimi K2.6 is Moonshot AI’s latest flagship open-source model, released on April 20, 2026 under a Modified MIT license. It is a native multimodal agentic model built on a 1-trillion parameter Mixture-of-Experts (MoE) architecture, with 32 billion parameters activated per token. The model is designed for long-horizon coding, autonomous execution, and multi-agent orchestration, and is […]</p>
NVIDIA Nemotron 3 Nano 30B API Benchmarks: Latency & Cost<p>About NVIDIA Nemotron 3 Nano 30B A3B NVIDIA Nemotron 3 Nano 30B A3B is a large language model trained from scratch by NVIDIA, designed as a unified model for both reasoning and non-reasoning tasks. It is part of the Nemotron 3 family — NVIDIA’s most efficient family of open models, built for agentic AI applications. […]</p>
Qwen3.5 9B API Benchmarks: Latency, Throughput & Cost<p>About Qwen3.5 9B Qwen3.5 9B is the flagship of Alibaba’s Qwen3.5 Small Model Series, released on March 2, 2026. It is a dense multimodal model combining Gated Delta Networks (a form of linear attention) with a sparse Mixture-of-Experts system, enabling higher throughput and lower latency during inference compared to traditional dense architectures. The architecture utilizes […]</p>
© 2026 DeepInfra. All rights reserved.