DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone and a context window of 1,048,576 tokens. It accepts text and image input and generates text autoregressively.
DeepInfra serves the model at fp8 with the full 1,048,576-token context, so the served window matches the model’s native window. The architecture targets input-heavy agentic workloads: the Causal Encoder-Decoder design activates 8B parameters per token during prefill and 16B during decode, which holds cost down on long prompts.
Model string: deepseek-ai/DeepSeek-V4.1-Flash
The release of DeepSeek-V4.1-Flash marks a significant leap forward in memory and compute efficiency. Unlike traditional dense models, DeepSeek-V4.1-Flash utilizes a Causal Encoder-Decoder (CED) architecture that activates only a fraction of its parameters during processing—specifically 8 billion parameters during the prefill phase and 16 billion during decoding.
These figures come from DeepSeek’s internal evaluation framework and describe the base model, not the instruct model served on DeepInfra. Scores within 0.3 of each other are treated as equivalent.
| Benchmark (Metric) | Shots | V4-Flash-Base | V4-Pro-Base | V4.1-Flash-Base |
|---|---|---|---|---|
| AGIEval (EM) | 3-5 | 83.9 | 84.4 | 83.4 |
| MMLU-Pro (EM) | 5 | 68.3 | 73.5 | 74.1 |
| SimpleQA-Verified (EM) | 25 | 30.1 | 55.2 | 42.3 |
| BBH (EM) | 3 | 86.9 | 87.5 | 86.1 |
| HumanEval (Pass@1) | 0 | 69.5 | 76.8 | 79.4 |
| GSM8K (EM) | 8 | 90.8 | 92.6 | 93.0 |
| MATH (EM) | 4 | 57.4 | 64.5 | 61.1 |
| DocVQA (LLM-Judge) | 4 | not reported | not reported | 95.6 |
V4.1-Flash-Base leads on MMLU-Pro, HumanEval and GSM8K. It trails both predecessors on AGIEval, BBH and MATH, and V4-Pro keeps a wide lead on SimpleQA-Verified. Backbone parameters are 284B for V4-Flash, 1.6T for V4-Pro and 552B for V4.1-Flash, so the comparison is not like for like on size.
All instruct results use reasoning_effort=100, temperature=1.0 and top_p=0.95. Code agent benchmarks run with the Minimal mode of DeepSeek Harness at a 1M-token context window, except DeepSWE v1.1, which uses the mini-SWE harness.
| Benchmark (Metric) | V4.1-Flash | V4-Pro | GPT-5.6 Sol | Opus-5.0 |
|---|---|---|---|---|
| Codeforces (Rating) | 3471 | 3348 | not reported | not reported |
| GPQA Diamond (Pass@1) | 90.9 | 92.4 | 94.1 | 93.4 |
| Terminal-Bench 2.1 (Pass@1) | 90.6 | 87.9 | 88.8 | 89.1 |
| Terminal-Bench 3.0 (Pass@1) | 30.0 | 11.8 | 34.4 | 43.3 |
| Terminal-Bench 4.0 (Pass@1) | 31.2 | 12.4 | 39.9 | 51.8 |
| DeepSWE v1.1 (Resolved) | 74.2 | 62.7 | 73.0 | 74.0 |
| AutomationBench (Pass@1) | 54.8 | 43.2 | 45.8 | 50.3 |
| CyberGym (Pass@1) | 88.1 | 83.3 | 84.5 | not reported |
V4.1-Flash takes the top score among these four on Codeforces, Terminal-Bench 2.1, DeepSWE v1.1, AutomationBench and CyberGym. Opus-5.0 and GPT-5.6 Sol stay well ahead on the harder Terminal-Bench 3.0 and 4.0 sets, and both lead on GPQA Diamond.
Scaffold choice moves the DeepSWE v1.1 number by roughly five points. The 74.2 figure uses mini-SWE. DeepSeek Harness Minimal returns 72.6, Claude Code 69.8 and Codex 65.6.
| Tier | Input | Output | Cached input |
|---|---|---|---|
| Standard | $0.20 | $0.60 | $0.006 |
| Priority (1.5x) | $0.30 | $0.90 | $0.009 |
| Flex (0.8x) | $0.16 | $0.48 | $0.0048 |
All figures are per 1M tokens.
Set service_tier to “priority” for faster time-to-first-token and higher throughput during peak demand, at a 50% surcharge. Set it to “flex” for a 20% discount in exchange for slower, best-effort scheduling: a flex request can wait up to ten minutes for capacity before it runs or returns an HTTP 429, so use it for work you can retry. Leave service_tier unset for standard real-time scheduling and pricing.
The response includes a service_tier field confirming which tier served the request. If a model does not support the tier you asked for, DeepInfra serves the request at standard tier and bills it at standard price without returning an error.
The reasoning_effort setting gives you a second cost lever, trading inference cost against accuracy.
The API is OpenAI-compatible. Point your existing OpenAI client at DeepInfra’s base URL and change the model name.
Base URL: https://api.deepinfra.com/v1/openai
Endpoint: /chat/completions
Auth header: Authorization: Bearer $DEEPINFRA_TOKEN
url "https://api.deepinfra.com/v1/openai/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $DEEPINFRA_TOKEN" \
-d '{
"model": "deepseek-ai/DeepSeek-V4.1-Flash",
"messages": [
{
"role": "user",
"content": "Explain the benefits of a Mixture-of-Experts (MoE) architecture."
}
],
"max_tokens": 150
}'Python and JavaScript both work by setting base_url to the DeepInfra endpoint and api_key to your DeepInfra token.
| Parameter | Notes |
|---|---|
| model | Required. deepseek-ai/DeepSeek-V4.1-Flash |
| messages | Required. Roles: system, user, assistant. Supports interleaved image content. |
| max_tokens | Maximum tokens to generate. See output limits below. |
| temperature | Sampling temperature, 0 to 2. Default 1.0. DeepSeek recommends 1.0 for this model. |
| top_p | Nucleus sampling threshold. Default 1.0. DeepSeek recommends 0.95 for this model. |
| stream | Set to true for server-sent events, terminating with [DONE]. |
| stop | Up to 4 sequences that halt generation. |
| n | Number of completion sequences. Default 1. |
| presence_penalty | Range -2.0 to 2.0. Default 0. |
| frequency_penalty | Range -2.0 to 2.0. Default 0. |
| response_format | Set to {“type”: “json_object”} for structured output. |
| tools, tool_choice | Function calling. |
| service_tier | “priority” or “flex”. Leave unset for standard. |
| fail_fast | Return HTTP 429 immediately instead of queueing when the model is at capacity. Default false. |
| reasoning_effort | Integer 1 to 100. Controls reasoning depth against cost. |
Maximum output tokens are model-dependent, with a hard cap of 16,384 tokens on most models. This is separate from the 1,048,576-token context window, which covers prompt and completion together.
For longer output, use response continuation: send a follow-up request with the previous response included as an assistant message, and the model picks up where it stopped. Continuation cannot push past the total context window, and a request that exceeds it returns a 400 error.
{
"id": "chatcmpl-guMTxWgpFf",
"object": "chat.completion",
"created": 1694623155,
"model": "deepseek-ai/DeepSeek-V4.1-Flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "..."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 15,
"completion_tokens": 16,
"total_tokens": 31,
"estimated_cost": 0.0000268
}
}The usage object returns an estimated_cost field alongside token counts, which is the fastest way to sanity-check spend during development.
DeepSeek-V4.1-Flash is optimized for cost-efficiency, offering tiered pricing to match different workload requirements.
Users can select a service_tier to balance cost and performance:
Note: The reasoning_effort setting allows you to further trade inference cost for accuracy. For the most up-to-date pricing, please visit the official DeepInfra pricing page.
DeepSeek-V4.1-Flash represents a paradigm shift in high-scale AI, combining a massive 552B parameter backbone with industry-leading KV cache compression. Whether you are building autonomous software engineering agents, processing million-token documents, or deploying multimodal applications, DeepSeek-V4.1-Flash provides the flexibility and power required at a fraction of the cost of traditional frontier models.
Next Steps:
DeepSeek V4.1 Flash API: Speed, Latency & Cost<p>DeepSeek V4.1 Flash (Reasoning, Max Effort) API Review Summary Metric Value Context Intelligence 40 (Artificial Analysis Intelligence Index) Well above median for comparable open-weight models (median: 18) Speed 211.5-545.6 output tokens/sec Notably fast; median: 68.9 t/s Latency (TTFT) 1.19s-5.28s (varies by provider) Competitive; median: 2.32s Price (DeepSeek API) $0.30/1M input, $1.20/1M output (peak) Cache discount: […]</p>
Fork of Text Generation Inference.The text generation inference open source project by huggingface looked like a promising
framework for serving large language models (LLM). However, huggingface announced that they
will change the license of code with version v1.0.0. While the previous license Apache 2.0
was permissive, the new on...
Fine-Tuning vs RAG vs Prompting: 2026 Guide<p>When an AI system yields unreliable answers, the root cause could be an unclear system prompt, missing context, poor retrieval quality, or simply using the wrong base model. Teams end up spending weeks experimenting with prompt changes, retrieval-augmented generation (RAG), or fine-tuning to improve response quality. But before deciding which technique to adopt, it is […]</p>
© 2026 DeepInfra. All rights reserved.