DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

XiaomiMiMo/
$0.39
in
$1.17
out
$0.078
cached
/ 1M tokens
$1.00
in
$3.00
out
$0.20
cached
/ 1M tokens
| Tier | Input | Output | Cached input |
|---|---|---|---|
Priority (1.5×)Learn More | $0.585 | $1.755 | $0.117 |
Flex (0.8×)Learn More | $0.312 | $0.936 | $0.0624 |
per 1M tokens — promotional pricing, 61% off already applied
MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).

b300
2026-05-03T02:03:53+00:00
GKZPammL
2026-05-03T02:03:53+00:00
© 2026 DeepInfra. All rights reserved.