DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

deepseek-ai logo

deepseek-ai/

DeepSeek-V4.1-Flash-Turbo

$0.30

in

$1.20

out

$0.006

cached

/ 1M tokens

TierInputOutputCached input
Priority (1.5×)Learn More
$0.45$1.80$0.009

per 1M tokens

DeepSeek-V4.1-Flash-Turbo is a speed-optimized deployment of DeepSeek-V4.1-Flash: the same FP8 weights, served with tensor parallelism and DSpark speculative decoding for higher per-request output speed and lower latency. Multimodal MoE with up to 1M tokens of context.

Deploy Private Endpoint
Supports Priority Tier
Public
Zero retention
fp8
1,048,576
JSON
Function
Multimodal
deepseek-ai/DeepSeek-V4.1-Flash-Turbo cover image
api
deepseek-ai/DeepSeek-V4.1-Flash-Turbo cover image
DeepSeek V4.1 Flash Turbo

Ask me anything

0.00s

You need to log in to use this model

Log In

Settings