DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Moonshot AI’s Kimi K3 represents a significant leap in open-weight AI model development. Released on July 16, 2026, this 2.8-trillion-parameter reasoning model has quickly become a focal point for developers seeking frontier-level intelligence with the flexibility of open weights. This analysis evaluates Kimi K3’s technical specifications, benchmark performance, and compares the leading API providers offering access to the model.
| Metric | Value | Context |
|---|---|---|
| Intelligence Index Score | 57 | Well above median 32 for open-weight models |
| Context Window | 1M tokens | ~1,573 A4 pages equivalent |
| Input Modalities | Text + Image | Multimodal reasoning model |
| Architecture | MoE (2.8T total / 104B active) | 16 of 896 experts activated per token |
| Input Pricing | $3.00/M tokens | Median: $1.75 |
| Output Pricing | $15.00/M tokens | Median: $10.00 |
| Cache Hit Price | $0.30/M tokens | 90% discount |
| Output Speed | 32.0 t/s | Below median for reasoning models of this class |
| TTFT Latency | 144.34s | High relative to median due to always-on max reasoning |
| Output Tokens (Intelligence Index) | 130M | Median: 63M |
| Evaluation Cost | $2,437.41 | Full Intelligence Index run |
Kimi K3 is Moonshot AI’s flagship open-weight model, representing the first publicly available AI system in the 3-trillion-parameter class. The model utilizes a Mixture of Experts (MoE) architecture with 2.8 trillion total parameters, activating 104 billion parameters during inference through its Stable LatentMoE framework (16 of 896 experts per token).
The architecture incorporates two key innovations: Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals (AttnRes), both designed to improve information flow through longer sequences and deeper model layers. Moonshot claims these structural advances provide approximately 2.5x the scaling efficiency compared to K2.
Kimi K3 is a reasoning model that uses extended chain-of-thought thinking to work through complex problems. At launch, it operates at maximum thinking effort by default, with low- and high-effort modes planned for subsequent updates. The model supports multimodal input (text and images) and generates text output.
| Specification | Value |
|---|---|
| Release Date | July 16, 2026 |
| Developer | Moonshot AI |
| Total Parameters | 2.8 trillion |
| Active Parameters | 104 billion |
| Architecture | MoE with Stable LatentMoE |
| Expert Configuration | 16 of 896 experts active |
| Context Window | 1,048,576 tokens (1M) |
| Input Modalities | Text, Image |
| Output Modalities | Text |
| Output Speed | 32.0 tokens/sec |
| Time to First Token | 144.34s |
| License | Kimi K3 License (commercial use requires separate agreement) |
| Weights Availability | Open weights (released July 27, 2026) |
Kimi K3 is available through multiple API providers, each offering distinct advantages for different deployment scenarios. The following analysis compares the primary options available to developers.
Moonshot’s first-party API serves as the reference implementation for Kimi K3. All benchmark figures in this analysis derive from Kimi’s API performance.
The first-party API provides the official baseline behavior used in referenced benchmarks and is recommended when exact alignment with Moonshot’s specifications is required.
Fireworks AI launched Kimi K3 support on day zero of the open-weights release. The platform has developed custom optimizations specifically for K3’s architecture.
Fireworks positions K3 as delivering performance comparable to closed frontier models at a meaningfully better cost-per-task. The platform is particularly suited for teams requiring immediate training capabilities alongside inference.
Together AI provides reliable API endpoints optimized for open-weights models, with full support for K3’s extensive context window.
Together AI is recommended for applications requiring stable generation quality with preserved reasoning content across conversation turns.
OpenRouter aggregates multiple Kimi K3 providers, offering flexible routing options for developers seeking optimal price-performance balance.
OpenRouter is ideal for developers already using its API ecosystem who want automatic provider selection based on real-time performance metrics.
| Provider | Best For | Key Advantage | Considerations |
|---|---|---|---|
| Kimi (First-Party) | Official baseline | Reference implementation | Higher pricing, slower performance |
| Fireworks | Training + Inference | Custom K3 kernels, US-hosted, ZDR | Day-zero support, LoRA training |
| Together AI | Long-horizon work | Preserved thinking history | Serverless and dedicated options |
| OpenRouter | Multi-provider routing | Automatic optimization | Aggregated pricing |
Kimi K3 achieves a score of 57 on the Artificial Analysis Intelligence Index, a composite benchmark evaluating models across reasoning, knowledge, mathematics, and coding. This places K3 well above the median score of 32 for comparable open-weight models.
Kimi K3’s pricing positions it in the mid-to-premium tier for open-weight models:
| Metric | Kimi K3 | Open-Weight Median |
|---|---|---|
| Input (per 1M) | $3.00 | $0.43 |
| Output (per 1M) | $15.00 | $1.20 |
| Cache Hit (per 1M) | $0.30 | — |
Kimi K3 represents a compelling option for developers requiring frontier-level reasoning capabilities with the flexibility of open weights. Its massive 1M-token context window, multimodal input support, and strong benchmark performance make it particularly suited for complex, long-horizon tasks in coding, research, and enterprise workflows.
The model’s primary limitations, high latency and slower output speed relative to peers, make it better suited for asynchronous, high-complexity tasks where reasoning accuracy is prioritized over immediate response times. For latency-sensitive applications, consider routing simpler queries to faster models while reserving K3 for complex reasoning tasks.
Kimi K3 is coming soon to DeepInfra. In the meantime, the previous generation, Kimi K2.5, is live on the platform, and the broader DeepInfra model catalog is worth monitoring for when K3 support goes live.
Art That Talks Back: A Hands-On Tutorial on Talking ImagesTurn any image into a talking masterpiece with this step-by-step guide using DeepInfra’s GenAI models.
Accelerating Reasoning Workflows with Nemotron 3 Nano on DeepInfraDeepInfra is an official launch partner for NVIDIA Nemotron 3 Nano, the newest open reasoning model in the Nemotron family. Our goal is to give developers, researchers, and teams the fastest and simplest path to using Nemotron 3 Nano from day one.
Kimi K2.5 API Benchmarks: Latency, Throughput & Cost<p>About Kimi K2.5 Kimi K2.5 is Moonshot AI’s flagship open-source reasoning model, released in January 2026. It is a native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens. The model features a Mixture-of-Experts (MoE) architecture with 1 trillion total parameters and 32 billion activated parameters. Kimi K2.5 […]</p>
© 2026 DeepInfra. All rights reserved.