DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
inclusionAI/
$0.06
in
$0.18
out
$0.012
cached
/ 1M tokens
The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. It’s mainly designed for multimodal agentic workflows, long-context understanding, and multi-step reasoning.

© 2026 DeepInfra. All rights reserved.