DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

GLM-5.3 is Z.ai’s reasoning model for long-horizon coding and agentic tasks, with a 1M-token context window. Serving it at scale raises practical questions about cost per token, output speed, and how long a request waits before the first answer token arrives.
This guide compares nine providers that offer GLM-5.3. Speed and latency figures come from Artificial Analysis benchmarks on a 10,000-token prompt, and pricing and features come from each provider’s published documentation. Because GLM-5.3 always reasons before answering, time to first answer token includes thinking time, so it runs well above the time to first chunk.
A quick reference by priority:
| Provider | Best For |
|---|---|
| DeepInfra | Lowest blended price of the 22 benchmarked providers, plus JSON mode, function calling, and private endpoints. |
| Z.ai | Direct first-party access and GLM Coding Plan subscriptions. |
| Fireworks AI | Fast output speed with a short wait to the first answer token. |
| GMI Cloud | Teams that want serverless inference and bare-metal GPU rental on one platform. |
| FriendliAI | High-throughput agent workloads backed by a published 99.99% uptime SLA. |
| Together AI | Serverless and dedicated deployments with a 99.9% SLA and three reasoning effort levels. |
| Telnyx | Teams already on Telnyx’s network that want inference on the same infrastructure. |
| Inco | The fastest output speed and lowest latency, at the highest blended price. |
| DigitalOcean | Teams already on DigitalOcean that want GLM-5.3 beside 70+ other models. |
DeepInfra is the best overall pick when cost per task drives the decision. Artificial Analysis lists it at $0.72 per 1M tokens blended, the lowest of the 22 providers it tracks for GLM-5.3.
The cached input rate matters most for agentic and RAG workloads, which resend large prompt prefixes on every turn. At $0.125 per 1M tokens, cached input costs about 78% less than the promotional input rate. The full rate card, including Flex, is on DeepInfra’s pricing page, and the GLM-5.3 API reference covers reasoning_effort and the OpenAI-compatible endpoint. Measured output speed sits mid-field, so applications that need the fastest streaming should weigh Inco (FAST) or Fireworks AI.
Z.ai is the model’s creator and offers first-party API access. The same API underpins its GLM Coding Plan subscriptions.
Z.ai ships new releases first and supports the coding plans directly. Its per-token rates run above DeepInfra’s current listing. GLM-5.3 shares its base model with GLM-5.2, which DeepInfra covers in its GLM-5.2 model overview.
Fireworks AI ranks fourth for output speed in Artificial Analysis testing and offers a shorter wait to the first answer token than most providers.
Fireworks pairs high throughput with one of the shorter waits to the first answer token, which suits interactive applications. Inco and Nebius post higher raw throughput.
GMI Cloud combines serverless inference with bare-metal GPU rental on one platform.
GMI suits teams that train or fine-tune on rented GPUs and want to serve the result from the same platform. Its measured GLM-5.3 throughput sits near the bottom of the field, second-lowest of 22 providers.
FriendliAI serves GLM-5.3 through its Model APIs, with dedicated endpoints available for reserved capacity.
FriendliAI publishes its uptime and speed claims platform-wide, so validate them on your own GLM-5.3 workload. The Artificial Analysis figures above are the independent reference point.
Together AI offers GLM-5.3 on serverless and dedicated infrastructure.
Together AI charges the same per-token rates as Z.ai’s first-party API. Its SLA and dedicated option suit production agents that need contractual uptime.
Telnyx runs GLM-5.3 on GPUs it owns and operates, reachable through an OpenAI-compatible Chat Completions API.
Telnyx suits teams already using its voice or messaging network, since inference stays on the same infrastructure. Independent speed data is not yet available.
Inco leads Artificial Analysis’s speed and latency rankings for GLM-5.3 through its Inco (FAST) tier.
Inco (FAST) suits real-time applications where response speed outweighs cost. Budget-sensitive workloads will find it the most expensive option.
DigitalOcean offers GLM-5.3 through Serverless Inference, alongside more than 70 other models on one endpoint.
DigitalOcean fits teams already running on its cloud who want GLM-5.3 next to other models under one bill. Latency-sensitive workloads should look at faster providers.
The right GLM-5.3 provider depends on which constraint binds first: price, speed, or platform fit. Measure latency as time to first answer token, since thinking time dominates the wait.
DeepInfra is the best overall pick when cost per task decides the choice. It pairs the lowest blended price in Artificial Analysis’s ranking with JSON mode, function calling, cached input at $0.125 per 1M tokens, and private endpoints. Workloads that need the fastest streaming should benchmark Inco (FAST) or Fireworks AI first.
Lighter workloads can run on GLM-5.3-Flash at a lower per-token price. Teams comparing generations can review DeepInfra’s GLM-5.2 pricing and cost comparison, and qualifying startups can apply to DeepStart for up to 1 billion free tokens.
DeepSeek-V4.1-Flash: Model Overview & Integration<p>DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone and a context window of 1,048,576 tokens. It accepts text and image input and generates text autoregressively. DeepInfra serves the model at fp8 with the full 1,048,576-token context, so the served window matches the model’s native window. The architecture targets input-heavy agentic workloads: the Causal […]</p>
Chat with books using DeepInfra and LlamaIndexAs DeepInfra, we are excited to announce our integration with LlamaIndex.
LlamaIndex is a powerful library that allows you to index and search documents
using various language models and embeddings. In this blog post, we will show
you how to chat with books using DeepInfra and LlamaIndex.
We will ...
GLM-5 API Benchmarks: Latency, Throughput & Cost<p>GLM-5 is the latest open-weights reasoning model released by Z AI (Zhipu AI) in February 2026, characterized by high “thinking token” usage. It is a Mixture of Experts (MoE) model with 744B total parameters and 40B active parameters, scaling up from GLM-4.5’s 355B parameters. The model was pre-trained on 28.5T tokens and features a 200K+ […]</p>
© 2026 DeepInfra. All rights reserved.