DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
deepseek-ai/
$0.30
in
$1.20
out
$0.006
cached
/ 1M tokens
| Tier | Input | Output | Cached input |
|---|---|---|---|
Priority (1.5×)Learn More | $0.45 | $1.80 | $0.009 |
per 1M tokens
DeepSeek-V4.1-Flash-Turbo is a speed-optimized deployment of DeepSeek-V4.1-Flash: the same FP8 weights, served with tensor parallelism and DSpark speculative decoding for higher per-request output speed and lower latency. Multimodal MoE with up to 1M tokens of context.
Ask me anything
You need to log in to use this model
Log InSettings
© 2026 DeepInfra. All rights reserved.