DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Kimi K3, developed by Moonshot AI, represents a landmark achievement in open-source artificial intelligence. As a 2.8-trillion-parameter native multimodal Mixture-of-Experts (MoE) model, Kimi K3 is engineered to handle demanding computational tasks, from complex software engineering and long-horizon agentic workflows to deep scientific research.
By combining a one-million-token context window with its architectural innovations, Kimi K3 offers frontier-level performance that rivals the world’s leading proprietary models. This overview covers the model’s capabilities, performance benchmarks, and how to integrate it via the DeepInfra platform.
Kimi K3 introduces several breakthroughs designed to maximize reasoning power while maintaining operational efficiency. Unlike traditional dense models, Kimi K3 utilizes a Stable LatentMoE framework. While the model contains 896 experts, it activates only 16 per token, totaling 104 billion active parameters. This design results in a 2.5x improvement in scaling efficiency compared to its predecessor, Kimi K2.
Kimi K3 consistently delivers top-tier results across reasoning, coding, and vision benchmarks, often setting the standard for open-weight models and matching proprietary systems like GPT-5.6 Sol and Claude Fable 5.
| Category | Benchmark | Kimi K3 (Max) | Claude Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Reasoning | GPQA Diamond | 93.5 | 92.6 | 94.1 |
| Coding | SWE-Marathon | 42.0 | 35.0 | 39.0 |
| Agentic | BrowseComp | 91.2 | 88.0 | 90.4 |
| Vision | Video-MME (w/ sub) | 90.0 | — | 89.5 |
| Vision | MMMU-Pro | 81.6 / 83.4 | 81.2 / 86.5 | 83.0 / 84.6 |
DeepInfra provides an OpenAI-compatible interface for Kimi K3, making it easy to integrate into existing workflows. The model is listed on the DeepInfra model catalog alongside DeepInfra’s broader lineup of open-weight models.
To interact with the API, you must use a DeepInfra API key. Retrieve your key from your DeepInfra Dashboard and pass it as a Bearer token in your HTTP headers:
Authorization: Bearer <YOUR_DEEPINFRA_API_KEY>
curl https://api.deepinfra.com/v1/openai/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $DEEPINFRA_API_KEY" \
-d '{
"model": "moonshotai/Kimi-K3",
"messages": [
{
"role": "user",
"content": "Explain the benefits of Kimi Delta Attention for long-context reasoning."
}
],
"temperature": 0.2
}'import os
import requests
api_key = os.getenv("DEEPINFRA_API_KEY")
url = "https://api.deepinfra.com/v1/openai/chat/completions"
headers = {
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json"
}
payload = {
"model": "moonshotai/Kimi-K3",
"messages": [{"role": "user", "content": "Write a Python function to optimize a GPU kernel."}],
"temperature": 0.2
}
response = requests.post(url, headers=headers, json=payload)
print(response.json())When sending requests to the /chat/completions endpoint, you can fine-tune the model’s behavior using the following parameters. The full parameter reference is available in the Kimi K3 API documentation.
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | Use moonshotai/Kimi-K3. |
| messages | array | Yes | Supports text and image inputs for multimodal tasks. |
| temperature | number | No | Controls randomness (e.g., 0.2 for deterministic outputs). |
| response_format | object | No | Use {“type”: “json_object”} for structured data. |
| tools | array | No | Define functions for agentic tool-calling workflows. |
| max_tokens | integer | No | Limits the length of the generated response. |
DeepInfra offers a transparent, usage-based pricing model for Kimi K3. The model supports cached input pricing, which significantly reduces costs for repeated long-context tasks by billing cached tokens at a fraction of the standard input rate.
| Token Type | Price per 1M Tokens |
|---|---|
| Input Tokens | $2.85 |
| Output Tokens | $14.25 |
| Cached Input Tokens | $0.285 |
For the most up-to-date information regarding volume discounts or tier-based pricing, refer to the official DeepInfra pricing page.
Kimi K3 stands as a powerful, cost-effective alternative to proprietary frontier models. Its combination of a 2.8T-parameter architecture, 1M-token context window, and native multimodality makes it a strong choice for developers building the next generation of AI agents and complex engineering tools. Here are the key takeaways:
To begin using Kimi K3 for your projects, visit the DeepInfra Dashboard to retrieve your API key or deploy a private endpoint for dedicated capacity.
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open-Weight AI Model Comparison<p>In the span of three months, three Chinese AI labs shipped open-weight models that individually would have rewritten the frontier story. Together, they signal something more structural: the open-weight tier is no longer a budget alternative to closed models. Kimi K3 (Moonshot AI, July 2026 — now also available through DeepInfra), DeepSeek V4 Pro (DeepSeek, […]</p>
GLM 5.2 vs Claude Opus 4.8: Pricing the Task, Not the Token<p>Every GLM 5.2 vs Claude Opus 4.8 comparison lands in the same place. Opus wins most coding benchmarks, GLM costs a fraction as much, pick according to your budget. That framing takes the price cards at face value, but it’s misleading. Price a finished unit of work instead of a million tokens and the gap […]</p>
Top 6 GLM-5.2 Max API Providers Compared<p>Deploying the GLM-5.2 (max) Mixture-of-Experts model — 753B total parameters with roughly 40B active per token and a 1M context window — requires infrastructure that separates production-grade API providers from the rest. This guide breaks down the top providers by throughput, latency, pricing, and quantization architecture. GLM-5.2 (max) API Review Summary (2026-06-27) TL;DR: Best Providers […]</p>
© 2026 DeepInfra. All rights reserved.