DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
nvidia/
$0.50
in
$2.20
out
$0.10
cached
/ 1M tokens
| Tier | Input | Output | Cached input |
|---|---|---|---|
Priority (1.5×)Learn More | $0.75 | $3.30 | $0.15 |
Flex (0.8×)Learn More | $0.40 | $1.76 | $0.08 |
per 1M tokens
Prompt cache retentionLearn More | Cache write |
|---|---|
Retain for 5m (1.25×) | $0.625 |
Retain for 1h (2×) | $1.00 |
per 1M tokens
Retained in whole blocks of 8,192 tokens; the remainder is billed as standard input. Reuse within the window bills at the cached input rate.
**Shown at the standard tier. Priority and Flex scale these rates the same way they scale input and output.
Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

pytrtllm2b300
2026-05-16T00:11:05+00:00
9V5zdJT3
2026-05-16T00:11:05+00:00
© 2026 DeepInfra. All rights reserved.