DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
thinkingmachines/
$0.45
in
$1.20
out
$0.10
cached
/ 1M tokens
Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active, trained on NVIDIA GB300 NVL72 systems. Like Inkling, it features native reasoning over audio and images, variable thinking effort

© 2026 DeepInfra. All rights reserved.