DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
moonshotai/
$0.68
in
$3.40
out
$0.136
cached
/ 1M tokens
| Tier | Input | Output | Cached input |
|---|---|---|---|
Priority (1.5×)Learn More | $1.02 | $5.10 | $0.204 |
Flex (0.8×)Learn More | $0.544 | $2.72 | $0.1088 |
per 1M tokens
Prompt cache retentionLearn More | Cache write |
|---|---|
Retain for 5m (1.25×) | $0.85 |
Retain for 1h (2×) | $1.36 |
per 1M tokens
Retained in whole blocks of 1,024 tokens; the remainder is billed as standard input. Reuse within the window bills at the cached input rate.
**Shown at the standard tier. Priority and Flex scale these rates the same way they scale input and output.
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

XY6BgGD6
2026-06-15T23:40:40+00:00
© 2026 DeepInfra. All rights reserved.