DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

zai-org/
$0.488
in
$1.56
out
$0.091
cached
/ 1M tokens
$0.75
in
$2.40
out
$0.14
cached
/ 1M tokens
| Tier | Input | Output | Cached input |
|---|---|---|---|
Priority (1.5×)Learn More | $0.7313 | $2.34 | $0.1365 |
Flex (0.8×)Learn More | $0.39 | $1.248 | $0.0728 |
per 1M tokens — promotional pricing, 35% off already applied
Prompt cache retentionLearn More | Cache write |
|---|---|
Retain for 5m (1.25×) | $0.6094 |
Retain for 1h (2×) | $0.975 |
per 1M tokens
Retained in whole blocks of 1,024 tokens; the remainder is billed as standard input. Reuse within the window bills at the cached input rate.
**Shown at the standard tier. Priority and Flex scale these rates the same way they scale input and output.
GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

© 2026 DeepInfra. All rights reserved.