DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

zai-org/
$0.15
in
$0.50
out
$0.03
cached
/ 1M tokens
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

© 2026 DeepInfra. All rights reserved.