We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

zai-org logo

zai-org/

GLM-5.3-Flash

$0.15

in

$0.50

out

$0.03

cached

/ 1M tokens

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

Deploy Private Endpoint
Public
fp8
1,048,576
JSON
Function
zai-org/GLM-5.3-Flash cover image
demoapi

XqKatdLz

2026-08-26T15:23:03+00:00