DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
inclusionAI/
$0.06
in
$0.18
out
$0.012
cached
/ 1M tokens
The multimodal version built on Ling-3.0-flash — 124B total / ~5.5B active per token, with native text, image, and video understanding. It’s mainly designed for multimodal agentic workflows, long-context understanding, and multi-step reasoning.

Ask me anything
You need to log in to use this model
Log InSettings
Ling-3.0-flash-VL, InclusionAI's next-generation native multimodal model. Built upon Ling-3.0-flash, it brings visual information into the complete process of understanding, reasoning, acting, and verification—advancing beyond image and video perception to solving real-world tasks through vision. With 124B total parameters, only 5.5B activated parameters per token, support for image and video inputs, and a context window of up to 1M tokens, Ling-3.0-flash-VL delivers powerful multimodal reasoning and agentic capabilities with exceptional efficiency.
Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens.
The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.
Overall, these designs make vision more than just an input, integrating it into the complete process of understanding, reasoning, planning, acting, and verification.

Ling-3.0-flash-VL achieves a score of 42 on the Artificial Analysis Intelligence Index v4.1.1, improving by 4 points over Ling-3.0-flash’s score of 38. The results show that extending the model with visual capabilities further improves its overall intelligence performance.

Across multimodal benchmarks, Ling-3.0-flash-VL demonstrates three distinct capability dimensions:

- Thinking mode is enabled by default. Unless otherwise specified, the default parameters for Ling-3.0-flash-VL are as follows:
temperature=0.6,top_p=0.95,top_k=20.- Terminal-Bench 2.1: Evaluated under the Artificial Analysis (AA) protocol using the default Terminus 2 harness, a unified 2-hour timeout, the provided JSON parser in preserve-thinking mode, and 3 runs per task (mean). Decoding uses temperature=1.0, max_new_tokens=32K, with a 256K context window.
© 2026 DeepInfra. All rights reserved.