DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepSeek-V4.1-Flash: Model Overview & Integration
Published on 2026.10.01 by DeepInfra
DeepSeek-V4.1-Flash: Model Overview & Integration

DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model with a 552B-parameter backbone and a context window of 1,048,576 tokens. It accepts text and image input and generates text autoregressively.

DeepInfra serves the model at fp8 with the full 1,048,576-token context, so the served window matches the model’s native window. The architecture targets input-heavy agentic workloads: the Causal Encoder-Decoder design activates 8B parameters per token during prefill and 16B during decode, which holds cost down on long prompts.

Model string: deepseek-ai/DeepSeek-V4.1-Flash

Architectural Breakthroughs and Latest News

The release of DeepSeek-V4.1-Flash marks a significant leap forward in memory and compute efficiency. Unlike traditional dense models, DeepSeek-V4.1-Flash utilizes a Causal Encoder-Decoder (CED) architecture that activates only a fraction of its parameters during processing—specifically 8 billion parameters during the prefill phase and 16 billion during decoding.

Key Innovations

  • Compressed Sparse Attention 2 (CSA2): Along with SWA Bounded Replay, this technology reduces the global KV cache footprint to just 890 bytes per token. This is a four-fold reduction compared to the previous V4-Flash and a 437-fold decrease from the original V1 model, effectively eliminating the memory bottleneck for long-context inference.
  • Controllable Reasoning: A standout feature of this release is the continuously controllable reasoning effort. Developers can now adjust a slider (1–100) to dynamically trade inference speed and cost for enhanced accuracy and reasoning depth.
  • Multimodal Integration: The model features DeepSeek-ViT, a custom-trained vision encoder that processes visual embeddings jointly with text from the start of pre-training, ensuring seamless multimodal understanding.

Benchmarks

Base model

These figures come from DeepSeek’s internal evaluation framework and describe the base model, not the instruct model served on DeepInfra. Scores within 0.3 of each other are treated as equivalent.

Benchmark (Metric)ShotsV4-Flash-BaseV4-Pro-BaseV4.1-Flash-Base
AGIEval (EM)3-583.984.483.4
MMLU-Pro (EM)568.373.574.1
SimpleQA-Verified (EM)2530.155.242.3
BBH (EM)386.987.586.1
HumanEval (Pass@1)069.576.879.4
GSM8K (EM)890.892.693.0
MATH (EM)457.464.561.1
DocVQA (LLM-Judge)4not reportednot reported95.6

V4.1-Flash-Base leads on MMLU-Pro, HumanEval and GSM8K. It trails both predecessors on AGIEval, BBH and MATH, and V4-Pro keeps a wide lead on SimpleQA-Verified. Backbone parameters are 284B for V4-Flash, 1.6T for V4-Pro and 552B for V4.1-Flash, so the comparison is not like for like on size.

Instruct model

All instruct results use reasoning_effort=100, temperature=1.0 and top_p=0.95. Code agent benchmarks run with the Minimal mode of DeepSeek Harness at a 1M-token context window, except DeepSWE v1.1, which uses the mini-SWE harness.

Benchmark (Metric)V4.1-FlashV4-ProGPT-5.6 SolOpus-5.0
Codeforces (Rating)34713348not reportednot reported
GPQA Diamond (Pass@1)90.992.494.193.4
Terminal-Bench 2.1 (Pass@1)90.687.988.889.1
Terminal-Bench 3.0 (Pass@1)30.011.834.443.3
Terminal-Bench 4.0 (Pass@1)31.212.439.951.8
DeepSWE v1.1 (Resolved)74.262.773.074.0
AutomationBench (Pass@1)54.843.245.850.3
CyberGym (Pass@1)88.183.384.5not reported

V4.1-Flash takes the top score among these four on Codeforces, Terminal-Bench 2.1, DeepSWE v1.1, AutomationBench and CyberGym. Opus-5.0 and GPT-5.6 Sol stay well ahead on the harder Terminal-Bench 3.0 and 4.0 sets, and both lead on GPQA Diamond.

Scaffold choice moves the DeepSWE v1.1 number by roughly five points. The 74.2 figure uses mini-SWE. DeepSeek Harness Minimal returns 72.6, Claude Code 69.8 and Codex 65.6.

Pricing

TierInputOutputCached input
Standard$0.20$0.60$0.006
Priority (1.5x)$0.30$0.90$0.009
Flex (0.8x)$0.16$0.48$0.0048

All figures are per 1M tokens.

Set service_tier to “priority” for faster time-to-first-token and higher throughput during peak demand, at a 50% surcharge. Set it to “flex” for a 20% discount in exchange for slower, best-effort scheduling: a flex request can wait up to ten minutes for capacity before it runs or returns an HTTP 429, so use it for work you can retry. Leave service_tier unset for standard real-time scheduling and pricing.

The response includes a service_tier field confirming which tier served the request. If a model does not support the tier you asked for, DeepInfra serves the request at standard tier and bills it at standard price without returning an error.

The reasoning_effort setting gives you a second cost lever, trading inference cost against accuracy.

Getting started

The API is OpenAI-compatible. Point your existing OpenAI client at DeepInfra’s base URL and change the model name.

Base URL: https://api.deepinfra.com/v1/openai

Endpoint: /chat/completions

Auth header: Authorization: Bearer $DEEPINFRA_TOKEN

url "https://api.deepinfra.com/v1/openai/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
  -d '{
      "model": "deepseek-ai/DeepSeek-V4.1-Flash",
      "messages": [
        {
          "role": "user",
          "content": "Explain the benefits of a Mixture-of-Experts (MoE) architecture."
        }
      ],
      "max_tokens": 150
    }'
copy

Python and JavaScript both work by setting base_url to the DeepInfra endpoint and api_key to your DeepInfra token.

Parameters

ParameterNotes
modelRequired. deepseek-ai/DeepSeek-V4.1-Flash
messagesRequired. Roles: system, user, assistant. Supports interleaved image content.
max_tokensMaximum tokens to generate. See output limits below.
temperatureSampling temperature, 0 to 2. Default 1.0. DeepSeek recommends 1.0 for this model.
top_pNucleus sampling threshold. Default 1.0. DeepSeek recommends 0.95 for this model.
streamSet to true for server-sent events, terminating with [DONE].
stopUp to 4 sequences that halt generation.
nNumber of completion sequences. Default 1.
presence_penaltyRange -2.0 to 2.0. Default 0.
frequency_penaltyRange -2.0 to 2.0. Default 0.
response_formatSet to {“type”: “json_object”} for structured output.
tools, tool_choiceFunction calling.
service_tier“priority” or “flex”. Leave unset for standard.
fail_fastReturn HTTP 429 immediately instead of queueing when the model is at capacity. Default false.
reasoning_effortInteger 1 to 100. Controls reasoning depth against cost.

Output limits

Maximum output tokens are model-dependent, with a hard cap of 16,384 tokens on most models. This is separate from the 1,048,576-token context window, which covers prompt and completion together.

For longer output, use response continuation: send a follow-up request with the previous response included as an assistant message, and the model picks up where it stopped. Continuation cannot push past the total context window, and a request that exceeds it returns a 400 error.

Response format

{
    "id": "chatcmpl-guMTxWgpFf",
    "object": "chat.completion",
    "created": 1694623155,
    "model": "deepseek-ai/DeepSeek-V4.1-Flash",
    "choices": [
        {
            "index": 0,
            "message": {
                "role": "assistant",
                "content": "..."
            },
            "finish_reason": "stop"
        }
    ],
    "usage": {
        "prompt_tokens": 15,
        "completion_tokens": 16,
        "total_tokens": 31,
        "estimated_cost": 0.0000268
    }
}
copy

The usage object returns an estimated_cost field alongside token counts, which is the fastest way to sanity-check spend during development.

Pricing and Service Tiers

DeepSeek-V4.1-Flash is optimized for cost-efficiency, offering tiered pricing to match different workload requirements.

Standard Pricing

  • Standard Input: $0.20 per 1 million tokens.
  • Cached Input: $0.006 per 1 million tokens.
  • Standard Output: $0.60 per 1 million tokens.

Service Tiers

Users can select a service_tier to balance cost and performance:

  • Priority Tier (1.5× cost): Optimized for low-latency production needs.
  • Flex Tier (0.8× cost): Ideal for non-urgent, cost-sensitive batch processing.

Note: The reasoning_effort setting allows you to further trade inference cost for accuracy. For the most up-to-date pricing, please visit the official DeepInfra pricing page.

Conclusion

DeepSeek-V4.1-Flash represents a paradigm shift in high-scale AI, combining a massive 552B parameter backbone with industry-leading KV cache compression. Whether you are building autonomous software engineering agents, processing million-token documents, or deploying multimodal applications, DeepSeek-V4.1-Flash provides the flexibility and power required at a fraction of the cost of traditional frontier models.

Next Steps:

  • Implement Prompt Caching: Use prompt_cache_key to lower costs for repetitive, long-context prompts.
  • Explore Multimodal: Integrate image analysis by passing image URLs in the messages array.
  • Fine-tune Reasoning: Experiment with the reasoning_effort parameter to find the perfect balance for your specific use case.
Related articles
DeepSeek V4.1 Flash API: Speed, Latency & CostDeepSeek V4.1 Flash API: Speed, Latency & Cost<p>DeepSeek V4.1 Flash (Reasoning, Max Effort) API Review Summary Metric Value Context Intelligence 40 (Artificial Analysis Intelligence Index) Well above median for comparable open-weight models (median: 18) Speed 211.5-545.6 output tokens/sec Notably fast; median: 68.9 t/s Latency (TTFT) 1.19s-5.28s (varies by provider) Competitive; median: 2.32s Price (DeepSeek API) $0.30/1M input, $1.20/1M output (peak) Cache discount: [&hellip;]</p>
Fork of Text Generation Inference.Fork of Text Generation Inference.The text generation inference open source project by huggingface looked like a promising framework for serving large language models (LLM). However, huggingface announced that they will change the license of code with version v1.0.0. While the previous license Apache 2.0 was permissive, the new on...
Fine-Tuning vs RAG vs Prompting: 2026 GuideFine-Tuning vs RAG vs Prompting: 2026 Guide<p>When an AI system yields unreliable answers, the root cause could be an unclear system prompt, missing context, poor retrieval quality, or simply using the wrong base model. Teams end up spending weeks experimenting with prompt changes, retrieval-augmented generation (RAG), or fine-tuning to improve response quality. But before deciding which technique to adopt, it is [&hellip;]</p>