DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepSeek’s new V4.1-Flash uses just 8 billion active parameters during input processing, out of a 552-billion-parameter backbone, and still beats the much larger DeepSeek-V4-Pro across every agentic benchmark DeepSeek published. Released in September 2026, it is built around a Causal Encoder-Decoder design that makes the gap between total and active parameters possible. For developers running input-heavy or long-context workloads, the practical effect is a lower cost per prefilled token without a drop in agentic capability.
It is now live on DeepInfra.
The headline architectural win is KV cache compression. V4.1-Flash’s global KV cache footprint lands at 890 bytes per token, roughly one quarter of its predecessor V4-Flash, which cuts the memory and storage costs that make long agentic runs expensive. On benchmarks that matter for real workloads, it earns a Codeforces rating of 3471, the highest in its comparison class, and scores 74.2 on DeepSWE v1.1, edging out Claude Opus 5.0 at 74.0. The model supports a 1,048,576-token context window with native multimodal input, runs at fp8 on DeepInfra, and is open-weights under an MIT license.
The most structurally notable thing about V4.1-Flash is its Causal Encoder-Decoder (CED) architecture, a real departure from the standard decoder-only transformer. The 40-layer backbone splits into a 20-layer causal encoder and a 20-layer decoder, and the decoder’s global KV cache is projected from the final encoder hidden states rather than built per layer. That asymmetry is what makes the active parameter counts possible: 8B active during prefill and 16B during decode, against a 552B backbone, alongside a 196B Engram conditional memory component that is accessed sparsely. For input-heavy workloads like agentic pipelines, the lower prefill cost is material.
The KV cache design deserves its own mention. V4.1-Flash uses Compressed Sparse Attention 2 (CSA2), which assigns each attention layer one of three static modes, Full, Reindex or Reuse, to share main KV and indexer K across layers. CSA2 together with FP4 main KV caching produces the 890-byte global footprint. A separate mechanism, SWA Bounded Replay, reconstructs sliding-window attention states by replaying recent tokens instead of persisting them to SSD, which cuts the persistent cache to roughly an eighth of V4-Flash. If you want to see how the numbers compare against current contemporaries rather than last generation’s, GLM-5.3 and Qwen3.8-Flash are the closest points of reference on DeepInfra today.
| Model | Global KV cache per token | Relative to V4.1-Flash |
|---|---|---|
| DeepSeek-V1 | approx. 389,000 bytes | 437x larger |
| DeepSeek-V4-Flash | approx. 3,560 bytes | 4x larger |
| DeepSeek-V4.1-Flash | 890 bytes | baseline |
DeepSeek publishes the 890-byte figure and the 4-fold and 437-fold ratios. The absolute byte counts for V1 and V4-Flash in the middle column are derived from those ratios rather than published directly. Persistent KV cache is a separate measurement: V4.1-Flash holds roughly an eighth of what V4-Flash needs on SSD.
For agents running long sessions with many cache hits, this translates directly into lower memory pressure and lower infrastructure cost.
On benchmarks, V4.1-Flash competes with, and in several agentic categories leads, models with far higher active parameter counts. Highlights from the instruct model evaluation:
The base model also hits 74.1 on MMLU-Pro against 73.5 for V4-Pro-Base, 79.4 on HumanEval and 93.0 on GSM8K, exceeding V4-Pro-Base on all three despite far fewer active parameters.
On capabilities, V4.1-Flash supports the full 1,048,576-token context window and is natively multimodal from pretraining, handling text and images through a DeepSeek-ViT trained from scratch with 2D-RoPE and a two-layer MLP projector. Vision benchmarks on the base model include 95.6 on DocVQA and 86.0 on RefCOCO-avg. The model also supports continuously controllable reasoning effort, an integer from 1 to 100, which lets you tune the cost and accuracy trade-off at inference time rather than choosing between a thinking and a non-thinking model.
Compared with what came before it: the architectural changes amount to a ground-up rebuild rather than a fine-tune or an incremental update. DeepSeek pre-trained the model from scratch on 45T tokens, rebuilt the post-training pipeline around large-scale automated agent task synthesis, and introduced the CED design. DeepInfra continues to host the earlier V4 models alongside it, so you can benchmark against them directly or follow release notes on the DeepInfra blog.
DeepSeek-V4.1-Flash is available now on DeepInfra under the model identifier deepseek-ai/DeepSeek-V4.1-Flash. Pricing is usage-based across three tiers: Standard at $0.20 per 1M input tokens and $0.60 per 1M output tokens, Priority at 1.5x for $0.30 and $0.90, and Flex at 0.8x for $0.16 and $0.48. Flex is a best-effort tier that can wait up to ten minutes for capacity before it runs or returns an HTTP 429. A rejected request never reaches the model, so nothing is billed for it.
Cached input tokens are priced at $0.006 per 1M on Standard, $0.009 on Priority and $0.0048 on Flex, a meaningful saving given that the architecture is built around input-heavy, cache-intensive workloads.
The model supports JSON mode, function calling and multimodal text-plus-image input out of the box. One practical note on limits: the 1,048,576-token context window covers prompt and completion together, and DeepInfra caps output at 16,384 tokens on most models. DeepSeek’s own recommendation of a 256K generation budget applies to running the weights yourself, not to the shared endpoint. If you need output past the cap, use response continuation by sending the previous response back as an assistant message.
DeepInfra gives you immediate API access with no infrastructure to manage: an OpenAI-compatible endpoint, usage-based billing, and Priority and Flex throughput tiers to match your latency requirements. The company operates a zero-retention policy, set out in its privacy policy, and holds SOC 2 and ISO 27001 certification. Live uptime is published on the DeepInfra status page, and the platform is backed by a $107M Series B raised to scale the inference cloud.
If you are evaluating alternatives alongside V4.1-Flash, NVIDIA Nemotron-3-Super-120B-A12B and Gemma 4 31B sit at different points on the price and capability curve. For embeddings, speech, reranking and image models, the other API surfaces follow the same billing model.
To make your first call, swap in your DeepInfra token and the model name:
curl "https://api.deepinfra.com/v1/openai/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $DEEPINFRA_TOKEN" \
-d '{
"model": "deepseek-ai/DeepSeek-V4.1-Flash",
"messages": [
{"role": "user", "content": "Hello!"}
]
}'
from openai import OpenAI
client = OpenAI(
api_key="$DEEPINFRA_TOKEN",
base_url="https://api.deepinfra.com/v1/openai",
)
response = client.chat.completions.create(
model="deepseek-ai/DeepSeek-V4.1-Flash",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
import OpenAI from "openai";
const openai = new OpenAI({
apiKey: "$DEEPINFRA_TOKEN",
baseURL: "https://api.deepinfra.com/v1/openai",
});
const response = await openai.chat.completions.create({
model: "deepseek-ai/DeepSeek-V4.1-Flash",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);The only things that change from any existing OpenAI-compatible integration are the base URL, https://api.deepinfra.com/v1/openai, your DeepInfra API key, and the model name.
DeepSeek-V4.1-Flash is a genuinely interesting engineering result: a model that cuts active parameter count and KV cache footprint sharply while moving benchmark scores in the right direction. The architecture choices are not cosmetic. The CED design and CSA2 compression have direct consequences for what it costs to run long agentic sessions at scale. For developers building code agents, autonomous pipelines, or anything that leans on large context and frequent cache hits, this belongs in your eval queue.
Grab an API key, run the curl call above, and compare it against whatever you are using today on your own workload rather than on a benchmark table. Early-stage teams building on DeepInfra can also apply to DeepStart for free inference credits.
NVIDIA Nemotron 3 Super: Model Overview & Integration Guide<p>The NVIDIA Nemotron 3 Super is a state-of-the-art 120-billion parameter hybrid Mixture-of-Experts (MoE) model designed to bridge the gap between high-compute efficiency and extreme accuracy. Engineered specifically for the next generation of AI development, Nemotron 3 Super excels in multi-agent applications, specialized agentic systems, and complex reasoning tasks. By utilizing a sophisticated architecture that activates […]</p>
OpenClaw Cost Optimization: Cut AI API Costs by 90%<p>A single ask in an OpenClaw session can cost more than a full evening of casual ChatGPT use. Ask your agent something simple, like which calendar event clashes with your flight, and the request that hits the API carries far more than your 12-token question. It also carries your SOUL.md, the tool schemas registered on […]</p>
Open vs Closed Source AI Models: Intelligence, Price & Speed Compared<p>The LLM landscape in 2026 looks nothing like it did two years ago. Back then the assumption was simple: if you wanted the best model, you paid OpenAI or Anthropic, and that was that. Open source models were a respectable second tier, good for experimentation, fine-tuning, and budget workloads, but not quite there for serious […]</p>
© 2026 DeepInfra. All rights reserved.