DeepInfra raises $107M Series B to scale the inference cloud — read the announcement
inclusionAI/
$0.06
in
$0.18
out
$0.012
cached
/ 1M tokens
The model prioritizes token efficiency and agentic inference at production scale, stretching what developers can achieve within limited token, latency, and serving-cost budgets.

1xB200-v1
2026-08-11T08:51:52+00:00
© 2026 DeepInfra. All rights reserved.