DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

zai-org logo

zai-org/

GLM-5.2

$0.563

in

$1.80

out

$0.105

cached

/ 1M tokens

$0.75

in

$2.40

out

$0.14

cached

/ 1M tokens

25% off
TierInputOutputCached input
Priority (1.5×)Learn More
$0.8438$2.70$0.1575
Flex (0.8×)Learn More
$0.45$1.44$0.084

per 1M tokens — promotional pricing, 25% off already applied

Prompt cache retentionLearn More
Cache write
Retain for 5m (1.25×)
$0.7031
Retain for 1h (2×)
$1.125

per 1M tokens

Retained in whole blocks of 1,024 tokens; the remainder is billed as standard input. Reuse within the window bills at the cached input rate.

**Shown at the standard tier. Priority and Flex scale these rates the same way they scale input and output.

GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

Deploy Private Endpoint
Supports Priority Tier
Supports Flex Tier
Public
Zero retention
fp4
1,048,576
JSON
Function
zai-org/GLM-5.2 cover image