DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
deepseek-ai/
$1.30
in
$2.60
out
$0.10
cached
/ 1M tokens
| Tier | Input | Output | Cached input |
|---|---|---|---|
Priority (1.5×)Learn More | $1.95 | $3.90 | $0.15 |
Flex (0.8×)Learn More | $1.04 | $2.08 | $0.08 |
per 1M tokens
Prompt cache retentionLearn More | Cache write |
|---|---|
Retain for 5m (1.25×) | $1.625 |
Retain for 1h (2×) | $2.60 |
per 1M tokens
Retained in whole blocks of 1,024 tokens; the remainder is billed as standard input. Reuse within the window bills at the cached input rate.
**Shown at the standard tier. Priority and Flex scale these rates the same way they scale input and output.
DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks.

g7gkjR6A
2026-04-24T07:37:18+00:00
© 2026 DeepInfra. All rights reserved.