DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

inclusionAI/

Ling-3.0-flash-VL

$0.06

in

$0.18

out

$0.012

cached

/ 1M tokens

The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. It’s mainly designed for multimodal agentic workflows, long-context understanding, and multi-step reasoning.

Deploy Private Endpoint
Public
fp16
131,072
JSON
Function
Multimodal
ProjectLicense
inclusionAI/Ling-3.0-flash-VL cover image