DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

zai-org logo

zai-org/

GLM-5.3-Flash

$0.075

in

$0.25

out

$0.015

cached

/ 1M tokens

$0.15

in

$0.50

out

$0.03

cached

/ 1M tokens

50% off

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Deploy Private Endpoint
Public
fp4
1,048,576
JSON
Function
Multimodal
zai-org/GLM-5.3-Flash cover image