DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

If you’re evaluating long-context reasoning models in late 2026, DeepSeek-V4.1-Flash is hard to ignore because the pricing is aggressive, the weights are open, and the provider market around it is already competitive. This is the rare model that shows up in both cost-sensitive buying conversations and serious agentic workloads: posted API pricing starts as low as $0.15 per million input tokens and $0.60 per million output tokens on some providers, while DeepInfra lists standard pricing at $0.20 input and $0.60 output with lower Flex pricing for teams optimizing harder on cost.
DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model from DeepSeek, released on September 10, 2026. It sits in the V4.1 family with a 552B-parameter backbone, activates 8B parameters on input and 16B on output, supports text-and-image input with text output, and ships with a 1M-token context window. It is also MIT licensed and available as open weights, which matters if you want optionality beyond a single hosted API. In the research sources, it is profiled both as a general public API model and as the reasoning variant at maximum effort, which is the context for many of the benchmark and pricing comparisons.
What makes it stand out is the combination of scale, openness, and deployment flexibility without obviously giving up on performance. Artificial Analysis scores it at 40 on its Intelligence Index, well above the reported median of 18, while OpenRouter cites a closely aligned 39.5 score and reports that it performs better than 82% of models compared. The model is also notably fast in third-party measurements: Artificial Analysis reports 211.5 output tokens per second and a 1.19s time to first token on DeepSeek’s API, while OpenRouter shows best-provider latency down to 0.42s and best throughput at 160 tokens per second across routed providers. For production teams, that combination matters more than abstract model positioning: you get a 1M-token window, reasoning support, tool calling, structured outputs, image understanding, and broad provider availability instead of being boxed into one deployment path.
For developers, ML engineers, and technical decision-makers, the real question is not whether the model is interesting. It is whether the economics hold up for your workload. DeepSeek-V4.1-Flash looks especially strong when you care about long prompts, caching, agent loops, or private deployment options. DeepInfra is particularly relevant there because it offers public access, private endpoints, zero data retention, JSON output, function calling, multimodal input, and three pricing tiers that make the tradeoff between cost and priority explicit. At the same time, the wider market matters too: OpenRouter exposes 19 providers with automatic failover, and some providers lead on throughput, latency, or tool-call reliability. That makes this a pricing guide worth reading closely rather than a model you price from a single vendor page.
DeepSeek-V4.1-Flash is best for teams that want an open-weight, 1M-context reasoning model without paying premium frontier-model rates. Pricing is fragmented in a useful way: DeepInfra starts at $0.16/$0.48 per million input/output tokens on Flex and $0.20/$0.60 on Standard, while OpenRouter shows provider pricing from $0.15/$0.60 up to $0.375/$1.50 across 19 providers; that makes provider choice part of the optimization, alongside competitors and alternatives surfaced in the research such as DeepSeek direct, OpenRouter routing, Baseten, SiliconFlow, Fireworks, NovitaAI, Together, Wafer, DigitalOcean, AtlasCloud, Morph, GMICloud, and benchmarked model peers including Opus-5.0, GPT-5.6 Sol, K3, GLM-5.3, DeepSeek-V4-Pro, and DeepSeek-V4-Flash.
| Best For | Provider Recommendation | Why |
| Lowest price / cost-sensitive workloads | DeepInfra Flex | Flex lists $0.16/M input, $0.48/M output, and $0.0048/M cached input, which is cheaper than DeepInfra Standard and competitive with the lowest posted market rates. |
| Proprietary or managed model access | DeepInfra Standard or Priority | DeepInfra offers public API access, private endpoint deployment, zero data retention, function calling, JSON output, multimodal input, and an explicit Priority tier for teams that want more managed delivery options. |
| Easiest onboarding / fastest time-to-first-call | DeepInfra Standard | DeepInfra exposes the model directly at deepseek-ai/DeepSeek-V4.1-Flash with standard hosted features, which is a simpler starting point than comparing nearly twenty routed providers. |
| RAG, document-heavy, or high-throughput use cases | DeepInfra Standard | The model supports a 1,048,576-token context window and DeepInfra pricing includes very cheap cached input at $0.006/M on Standard, which fits retrieval-heavy and repeated-context workloads well. |
| Lowest latency | Baseten via OpenRouter | OpenRouter reports Baseten at roughly 0.41s average latency and 1.16s end-to-end latency, making it the strongest choice when responsiveness matters most. |
| Highest throughput | SiliconFlow via OpenRouter | OpenRouter lists SiliconFlow as the top throughput provider at about 130 tokens per second on average, ahead of NovitaAI and DeepSeek. |
| Tool-calling reliability | Fireworks via OpenRouter | Fireworks shows the lowest average tool-call error rate in the provider table at 0.09%, which is valuable for agent workflows. |
| Structured output reliability | Morph via OpenRouter | Morph has the lowest reported structured output error rate at 3.05%, making it the cleanest pick for schema-constrained responses. |
| Maximum provider redundancy | OpenRouter auto-routing | OpenRouter serves the model through 19 providers and supports automatic failover, with 99.92% 24-hour availability reported with routing versus 93.22% without routing. |
DeepSeek-V4.1-Flash is cheap enough to look simple on a pricing page and expensive enough to punish lazy accounting in production. With a 1M-token context window, token categories matter more than usual.
| Token type | What it is | Why it matters |
| Input tokens | Tokens in the prompt you send to the model, including system messages, user messages, tool definitions, schemas, documents, and image-derived tokens | This is the base cost of every request. With a 1M-token context window, long prompts can dominate spend fast. |
| Output tokens | Tokens the model returns, including final answers and usually any verbose reasoning-style responses you ask it to produce | Output is priced higher than input on most providers for this model. Verbose answers, tool traces, and over-generous max token limits can spike cost. |
| Cached input tokens | Reused prompt tokens that the provider can serve from cache instead of recomputing from scratch | This is where long-context apps get cheaper. Good cache hit rates can materially change effective cost. |
| Reasoning tokens | Tokens consumed during the model’s internal reasoning workflow or exposed through provider-specific reasoning controls | These can increase total usage even if the visible answer is short. Agent and coding workflows feel this first. |
| Completion max tokens | Your configured upper bound for generated output | Not a separate billing category, but a dangerous one. Set it too high and the model has room to burn money. |
| Tool / schema overhead tokens | Tokens added by function definitions, JSON schemas, tool results, and structured-output wrappers | Small in toy examples, annoying at scale. Tool-heavy agent systems pay this tax on nearly every turn. |
A few cost patterns show up immediately:
Provider choice changes more than the sticker price. For this model, the main token-cost tradeoff is simple: some providers are better for low baseline cost, others are better when caching, routing, or reliability changes effective spend.
| Provider / route | Token cost advantage | Token cost disadvantage | Best fit |
| DeepInfra Flex | Lowest DeepInfra rates: $0.16/M input, $0.48/M output, $0.0048/M cached input. Very strong for repeated-context apps. | Flex tiers usually exist for a reason. If your workload is latency-sensitive or needs stricter priority handling, the cheaper token price may come with operational tradeoffs. | Batch jobs, async agents, large-scale RAG, cost-sensitive evaluation runs |
| DeepInfra Standard | Balanced pricing at $0.20/M input and $0.60/M output, plus very cheap cached input at $0.006/M. Easier to model than routed multi-provider setups. | Not the absolute lowest posted input price in market comparisons. If you only care about raw prompt cost on fresh requests, some OpenRouter providers list lower entry pricing. | Default production choice, especially for long prompts and reusable context |
| DeepInfra Priority | Predictable managed pricing and same token model as Standard, just with higher service priority. Useful when uptime and queue behavior matter more than shaving pennies. | Straight 1.5× premium: $0.30/M input, $0.90/M output, $0.009/M cached input. You feel that quickly on verbose agent workloads. | Business-critical flows where response priority matters more than lowest token cost |
| OpenRouter lowest-cost providers | Lowest posted input floor at $0.15/M and common output floor at $0.60/M. Good when you want market competition and don’t mind provider variation. | Cache read is often worse than DeepInfra’s rates. Also, the cheapest listed provider is not always the best end-to-end value once failover behavior, output verbosity, and tool retries show up. | Teams optimizing fresh-request cost and willing to benchmark providers |
| OpenRouter auto-routing | Can reduce the operational cost of outages and provider failures. Failed runs, retries, and fallback logic have token costs too, and routing can avoid some of that waste. | You lose some pricing determinism. Actual spend can reflect routed provider mix, and weighted average prices may differ from the lowest posted rates. | High-availability systems where reliability is part of cost control |
| High-priced OpenRouter providers | Sometimes justified by better latency, throughput, or tool-call reliability, which can lower total workflow cost in agent systems. | Up to $0.375/M input and $1.50/M output is a big jump. If your workload is mostly straightforward generation, you may just be paying extra for speed you do not need. | Interactive agents, time-sensitive tooling, workflows where retries are expensive |
If you want to run DeepSeek-V4.1-Flash hard without paying premium-provider tax, DeepInfra is the provider to look at first. Its bare-metal infrastructure matters because it cuts out a lot of the virtualization overhead that can drag on both performance and cost, which is exactly what power users care about when requests get large and frequent. DeepInfra is also typically 50–80% cheaper than major cloud competitors, so the value proposition is straightforward: more throughput headroom, lower unit economics, and less friction when you scale. That makes it an especially strong fit for developers, high-volume API users, and cost-conscious teams that want serious model access without bloated infrastructure pricing. If you want to size it up against the previous generation, our benchmark comparison of V4 Flash against Qwen3.6 and GLM-4.6 is a useful reference for how this family of models actually performs on reasoning tasks.
| Model Name | Best Use Case | Context Window | Input Price (per 1M tokens) | Output Price (per 1M tokens) |
| DeepSeek-V4.1-Flash (Flex) | Lowest-cost production and batch workloads | 1,048,576 tokens | $0.16 | $0.48 |
| DeepSeek-V4.1-Flash (Standard) | General production use with balanced cost and features | 1,048,576 tokens | $0.20 | $0.60 |
| DeepSeek-V4.1-Flash (Priority) | Latency- and priority-sensitive workloads | 1,048,576 tokens | $0.30 | $0.90 |
| DeepSeek V4.1 Flash (Reasoning, Max Effort) | Maximum-effort reasoning evaluation and advanced agent tasks | 1M tokens | $0.30 | $1.20 |
| DeepSeek V4.1 Flash via OpenRouter floor pricing | Lowest posted routed-provider entry pricing | 1,048,576 tokens | $0.15 | $0.60 |
Why This Matters: On DeepInfra, DeepSeek-V4.1-Flash starts at $0.16 per 1M input tokens and $0.48 per 1M output tokens on Flex, or $0.20/$0.60 on Standard. Even against the broader market reference points in the research, that keeps it near the floor on input cost while undercutting higher posted rates like $0.30/$1.20 for the max-effort reasoning pricing profile. If you expect large prompt volumes or lots of repeated calls, those per-million differences add up fast.
For teams planning real volume rather than one-off experiments, DeepInfra is the kind of provider that can materially improve the economics. If cost control is part of the architecture decision, it deserves a serious benchmark in your stack.
Below are practical developer scenarios where DeepInfra is a particularly strong way to run DeepSeek-V4.1-Flash: low token rates, very cheap cached input, and feature support like JSON output, function calling, multimodal input, zero retention, and private endpoints all map cleanly to real production patterns.
A developer team is building an internal support copilot that repeatedly sends the same system prompt, retrieval wrapper, and policy instructions, plus fresh user questions and retrieved passages. This is exactly the kind of workload where DeepInfra’s cached input pricing helps.
| Metric | Value |
| Volume | 50M input tokens + 10M output tokens per month |
| Model | DeepSeek-V4.1-Flash |
| Provider | DeepInfra Standard |
| Input Tokens | 50M |
| Output Tokens | 10M |
| Monthly Cost | $16.00 |
Cost math:
Comparison: The same workload on a higher-priced provider at $0.30/M input and $1.20/M output would cost $27.00, so DeepInfra Standard is $11.00 less per month on this usage.
A platform team runs nightly code review jobs across repositories, generating structured findings and suggested fixes. This is a classic cost-sensitive async workflow: lots of requests, plenty of output, and no need to pay a premium for priority handling.
| Metric | Value |
| Volume | 200M input tokens + 80M output tokens per month |
| Model | DeepSeek-V4.1-Flash |
| Provider | DeepInfra Flex |
| Input Tokens | 200M |
| Output Tokens | 80M |
| Monthly Cost | $70.40 |
Cost math:
Comparison: The same workload on a higher-priced provider at $0.375/M input and $1.50/M output would cost $195.00, so DeepInfra Flex is $124.60 less per month.
A developer builds an agent that reads issue tickets, calls internal tools, writes JSON summaries, and proposes code changes. These systems usually accumulate token overhead from tool definitions, schemas, and multi-turn loops, so stable pricing and full API feature support matter as much as raw model quality. If agentic capability is central to your use case, the DeepSeek-V4-Flash-0731 release is worth reviewing too, since it substantially enhanced agentic behavior over the preview version in the V4 line.
| Metric | Value |
| Volume | 120M input tokens + 40M output tokens per month |
| Model | DeepSeek-V4.1-Flash |
| Provider | DeepInfra Standard |
| Input Tokens | 120M |
| Output Tokens | 40M |
| Monthly Cost | $48.00 |
Cost math:
Comparison: The same workload on DeepInfra Priority at $0.30/M input and $0.90/M output would cost $72.00, so Standard saves $24.00 per month if you do not need priority handling.
A product team is processing screenshots, PDFs, and form images, then converting them into structured text output for downstream systems. Since DeepSeek-V4.1-Flash natively supports text-and-image input, this is a natural fit. Teams comparing multimodal options can also look at the experimental DeepSeek-V4-Flash-Vision-Exp model, which extends the V4-Flash architecture specifically for visual understanding.
| Metric | Value |
| Volume | 75M input tokens + 25M output tokens per month |
| Model | DeepSeek-V4.1-Flash |
| Provider | DeepInfra Standard |
| Input Tokens | 75M |
| Output Tokens | 25M |
| Monthly Cost | $30.00 |
Cost math:
Comparison: The same workload on a higher-priced provider at $0.30/M input and $1.20/M output would cost $52.50, making DeepInfra Standard $22.50 less per month.
A team has an agent framework that reuses a large shared prompt prefix across many turns: system instructions, tool specs, schemas, and persistent workspace context. This is one of the clearest cases for DeepInfra because cached input is priced very aggressively.
| Metric | Value |
| Volume | 100M cached input tokens + 20M output tokens per month |
| Model | DeepSeek-V4.1-Flash |
| Provider | DeepInfra Standard |
| Input Tokens | 100M cached input |
| Output Tokens | 20M |
| Monthly Cost | $12.60 |
Cost math:
Comparison: The same workload at an OpenRouter floor cache-read rate of $0.015/M and output rate of $0.60/M would cost $13.50, so DeepInfra Standard is cheaper by $0.90 per month before you factor in any other provider differences.
An enterprise engineering org wants an internal coding and docs assistant with zero data retention and the option to move to a private endpoint later, but it still wants to start on a public API first.
| Metric | Value |
| Volume | 30M input tokens + 15M output tokens per month |
| Model | DeepSeek-V4.1-Flash |
| Provider | DeepInfra Standard |
| Input Tokens | 30M |
| Output Tokens | 15M |
| Monthly Cost | $15.00 |
Cost math:
Comparison: The same workload on a higher-priced provider at $0.375/M input and $1.50/M output would cost $33.75, so DeepInfra Standard is $18.75 less per month.
For developers, the pattern across these scenarios is consistent: if you want DeepSeek-V4.1-Flash as a real production building block rather than a benchmark curiosity, DeepInfra is strongest when your workload includes repeated context, structured outputs, tool calls, multimodal input, or governance requirements like zero retention and private deployment options.
Choosing a provider for DeepSeek-V4.1-Flash is less about picking a winner and more about matching your workload’s actual cost drivers to the right pricing structure. The model is capable and broadly available — that part is settled. What changes your bill is whether your prompts are fresh or cached, whether your agents are verbose, and whether you need features like zero data retention or private endpoints baked into the same deployment.
The clearest decision criteria are token economics and feature coverage. If your workload involves repeated context — shared system prompts, retrieval wrappers, persistent agent state — DeepInfra’s cached input pricing at $0.006/M on Standard and $0.0048/M on Flex is a material advantage over providers that look cheaper on fresh input alone. If you need structured outputs, function calling, and multimodal support without stitching together multiple providers, DeepInfra covers all of it under one API surface. And if you are running high-volume batch jobs where priority handling is not a concern, DeepInfra’s Flex tier gives you the lowest posted rates without giving up production feature support.
For teams that want to evaluate the broader DeepSeek model family alongside V4.1-Flash, DeepInfra’s DeepSeek model catalog surfaces the full lineup — useful if you are benchmarking V4.1-Flash against V4-Flash or V4-Pro for a specific task before committing to one architecture. You can also browse the full DeepInfra model library if you want to compare against non-DeepSeek options in the same pricing framework. The pricing differences across that family are real, and running a quick cost projection on your actual prompt mix takes less time than it sounds.
If you are ready to run the model, the DeepSeek-V4.1-Flash API page on DeepInfra is the fastest place to start — you can test it in the demo, check the endpoint details, and move to a production integration without needing to talk to anyone first.
Data Sovereignty AI: Why Open-Weight Models Matter<p>The prototype worked. The demo landed. Then the security questionnaire shows up, and question 41 asks which legal entity can be compelled to produce the contents of your prompts. That one question stalls more AI rollouts than latency or accuracy ever have. Inside hospitals, banks, law firms, and defense subcontractors, data sovereignty AI requirements have […]</p>
Beat AI Subscription Fatigue With One API<p>Open your company card statement and scroll the recurring charges. Twenty dollars for a chat assistant, twenty more for a coding copilot, fifteen for an image API, another forty for the automation glue that wires them together. None of them is expensive on its own. Together they are a slow leak you stopped noticing months […]</p>
MiMo-V2.5 Model Documentation and Integration Guide<p>MiMo-V2.5 is a native omnimodal model developed by XiaomiMiMo, designed to process and understand text, image, video, and audio through a unified architecture rather than relying on “bolted-on” components for each modality. Built on a 310-billion-parameter Sparse Mixture of Experts (MoE) architecture — with only 15 billion parameters activated during inference — MiMo-V2.5 offers a […]</p>
© 2026 DeepInfra. All rights reserved.