DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepSeek V4 Pro is an open-weight Mixture-of-Experts model with 1.6T total parameters, 49B active parameters, and a 1M-token context window. It ships under the MIT license and supports JSON mode and function calling. The V4 Pro 0813 release is the GA version of the model first previewed on April 24, 2026.
DeepSeek V4 Pro costs $1.30 per 1M input tokens, $2.60 per 1M output tokens, and $0.10 per 1M cached tokens on DeepInfra. Those rates sit 25% below the previous DeepInfra pricing of $1.74 input and $3.48 output, and cached input is 31% cheaper than the earlier $0.145.
Artificial Analysis now tracks 10 API providers for the model. DeepInfra ranks second on blended price at $0.59 per 1M tokens, behind Systalyze at $0.33. Output speed ranges from 56 to 279.9 tokens per second across providers, so cost and latency requirements should drive your choice. This guide covers current pricing, provider benchmarks, service tiers, and five worked monthly cost scenarios.
For developers and ML teams, the practical question is not whether DeepSeek V4 Pro is interesting. It is whether you want the cheapest access path, the fastest output speed, the lowest raw time to first token, or a platform setup that fits your stack. The benchmark data here makes that evaluation unusually concrete: Fireworks leads on output speed and end-to-end latency, DeepInfra ties for the lowest benchmarked price while adding cached-token pricing and private deployment support, Together.ai posts the best raw time to first token at 0.99 seconds, and OpenRouter offers another access path with much lower listed per-token rates on its model page.
DeepSeek V4 Pro is an open-weight reasoning model for coding, long-context analysis, and agent workflows. Artificial Analysis benchmarks it across 10 API providers, and DeepInfra ranks second on blended price at $0.59 per 1M tokens. DeepInfra lists $1.30 input, $2.60 output, and $0.10 cached per 1M tokens, with optional Priority and Flex tiers. Choose DeepInfra when cost control, caching, and data retention matter most. Choose a faster provider when response time is the main constraint.
| Best For | Provider Recommendation | Why |
| Cost-sensitive production workloads | DeepInfra | Second-lowest blended price on Artificial Analysis at $0.59 per 1M tokens, with list rates of $1.30 input and $2.60 output. |
| Single API across many providers | OpenRouter | Lists 16 providers for the model and can reroute requests to another provider when one returns errors. |
| Lowest first-chunk latency | Baseten | Median first chunk arrives in 0.65 seconds on Artificial Analysis, ahead of DigitalOcean at 0.84 and DeepInfra at 1.23. |
| RAG and repeated-context workloads | DeepInfra | Cached input costs $0.10 per 1M tokens, a 92% discount to the $1.30 standard input rate. |
| Maximum output speed and lowest answer latency | Systalyze | Generates 279.9 tokens per second and returns the first answer token in 8.65 seconds, ahead of Crusoe and DeepSeek. |
| Lowest blended price | Systalyze | Posts a $0.33 blended price per 1M tokens on a 7:2:1 cache, input, and output mix, ahead of DeepInfra at $0.59. |
| Batch and non-urgent workloads | DeepInfra Flex | The Flex tier bills $1.04 input and $2.08 output per 1M tokens, a 20% discount, in exchange for best-effort scheduling. |
| Direct access from the model creator | DeepSeek | The first-party API lists $1.32 input and $3.96 output per 1M tokens and runs at 109.9 tokens per second. |
| Latency-sensitive workloads on DeepInfra | DeepInfra Priority | The Priority tier bills $1.95 input and $3.90 output per 1M tokens for faster first tokens and higher throughput at peak demand. |
| Data retention and compliance controls | DeepInfra | Zero retention, SOC 2 and ISO 27001 certification, and private endpoint deployment. |
Sources: Artificial Analysis (DeepSeek V4 Pro 0813, Max reasoning effort, 10,000-token workload), plus the DeepInfra and OpenRouter model pages, accessed October 2026.
Several details of DeepSeek V4 Pro pricing and availability have shifted since its launch. The table below summarizes the current picture.
| Area | Earlier version of this guide | Current |
| DeepInfra input price | $1.74 per 1M tokens | $1.30 per 1M tokens |
| DeepInfra output price | $3.48 per 1M tokens | $2.60 per 1M tokens |
| DeepInfra cached input price | $0.145 per 1M tokens | $0.10 per 1M tokens |
| Serving precision | FP4 | fp8, per the DeepInfra model page |
| Model release | April 24, 2026 preview | V4 Pro 0813 GA, released August 13, 2026 |
| DeepInfra service tiers | Standard only | Standard, Priority (1.5×), and Flex (0.8×) |
| Providers benchmarked | 6 | 10 |
| Blended price method | 3:1 input-to-output ratio | 7:2:1 cache, input, and output ratio |
| DeepInfra blended price | $2.17 per 1M tokens | $0.59 per 1M tokens |
| Cached-token pricing | Listed for DeepInfra only | Listed for most providers on OpenRouter and Artificial Analysis |
The two blended prices use different ratios, so they are not directly comparable. The current Artificial Analysis blend weights cached input at 70%, which is why blended figures now sit well below list prices.
If you have ever looked at a model bill and thought, “that prompt was not that long,” this is the part that explains why it was.
A token is a chunk of text the model reads or generates. It is not the same thing as a word. Short words may be one token. Longer words, code, JSON, whitespace patterns, and weird punctuation splits can turn into more. Pricing is based on tokens, so the shape of your workload matters more than the raw character count.
| Token type | What it is | Why it matters |
| Input tokens | Tokens you send in the request prompt | Usually the biggest driver for RAG, long-context, and agent workloads |
| Output tokens | Tokens the model generates back | Expensive when you ask for long answers, code diffs, or verbose JSON |
| Cached tokens | Reused prompt tokens billed at a reduced rate when supported | Can materially cut cost for repeated context-heavy requests |
| Blended tokens | A benchmarked combination of cache, input, and output token costs | Good for provider comparison, bad for forecasting if your traffic mix is different |
DeepSeek V4 is one of those models where the “headline price” can mislead you if you do not check how the provider bills input versus output.
| Provider | Token cost profile | Advantage | Disadvantage |
| DeepInfra | $1.30 input / $2.60 output / $0.10 cached per 1M | Second-lowest blended price in current provider data, with the largest cache discount at 92% off standard input | Output speed trails faster providers at under 60 tokens per second |
| DeepSeek (first-party) | $1.32 input / $3.96 output / $0.044 cached per 1M | Access directly from the model creator, with output speed over 100 tokens per second | Output price runs higher than DeepInfra and most third-party providers |
| Parasail | $0.45 input / $3.48 output / $0.10 cached per 1M | Lowest input price among currently tracked providers, with sub-second latency | Output price sits close to DeepInfra despite the lower input rate |
| DigitalOcean | $1.04 input / $2.09 output / $0.21 cached per 1M | Highest reported uptime among tracked providers at 99.91% | Cached input price runs about double DeepInfra’s rate |
| SiliconFlow | $1.50 input / $3.14 output / $0.14 cached per 1M | Specs mirror the official DeepSeek API, useful as a routing fallback | Higher latency than DeepInfra at just over 2 seconds to first token |
| Novita | $1.60 input / $3.20 output / $0.14 cached per 1M | Comparable throughput to DeepInfra, with full 1M context | Input and output prices both run higher than DeepInfra |
| Azure | $1.91 input / $3.83 output / $0.16 cached per 1M | Enterprise billing and compliance tooling already in place for Azure customers | Highest token pricing among the providers compared here |
| OpenRouter | Pricing varies by routed provider, from about $0.21 to $1.91 input per 1M | Automatic failover across providers if one returns an error | Cache and cost depend on which underlying provider handles the request |
Source: OpenRouter provider listing for deepseek/deepseek-v4-pro and Artificial Analysis, accessed October 2026.
DeepInfra and Parasail show the most useful cache discounts in the current provider data.
Base token rates now differ meaningfully across providers, unlike the earlier benchmark set.
Output tokens are where teams get sloppy and bills get weird.
OpenRouter pricing now tracks each underlying provider directly.
Cache-token pricing is no longer a DeepInfra-only feature.
Practical rule of thumb for DeepSeek V4 Pro billing
Practical rule of thumb for DeepSeek V4 billing
DeepSeek V4 Pro supports three reasoning modes: Non-think, Think High, and Think Max. Each trades response quality against the number of tokens the model generates before answering.
Non-think mode returns fast, direct responses with no visible reasoning step, suited to routine, low-risk tasks. Think High mode adds a visible reasoning pass for complex problem-solving and planning. Think Max pushes reasoning effort to its limit for the hardest tasks, at the cost of the most output tokens.
Reasoning tokens bill at the same output rate as the final answer. On DeepInfra, that means $2.60 per 1M tokens on standard pricing, so a Think Max request can cost meaningfully more than the same prompt run in Non-think mode. Artificial Analysis reports that DeepSeek V4 Pro 0813 at max reasoning effort generated 160 million output tokens across its Intelligence Index evaluation, above the median for comparable open-weight models.
If your app defaults to a high reasoning effort for every request, confirm that the task actually needs it. Routine classification, extraction, or short-answer tasks rarely benefit from Think Max and will cost more than necessary if it runs by default.
DeepSeek V4 Pro is the largest model in the family, built for workloads that need the strongest reasoning and coding results regardless of cost. DeepSeek V4 Flash trades some capability for lower cost and faster inference. DeepSeek V4.1 Flash is the newer multimodal release, built on a different architecture that activates far fewer parameters per token and adds native image understanding.
If your workload is reasoning-heavy, long-context, or code-focused and cost is a secondary concern, V4 Pro remains the stronger pick. If you need lower cost per token, faster throughput, or image input, V4.1 Flash is worth evaluating directly against V4 Pro for your specific task before committing. See the [DeepSeek V4.1 Flash pricing guide] for a full cost breakdown of that model.
DeepInfra is a strong choice for DeepSeek V4 Pro because it combines straightforward token pricing with infrastructure built for production deployment. The platform runs on bare-metal infrastructure, which removes a layer of virtualization overhead that can affect both performance consistency and cost. Standard pricing sits at $1.30 per 1M input tokens, $2.60 per 1M output tokens, and $0.10 per 1M cached tokens, with Priority and Flex tiers available for workloads that need faster or cheaper service. If you want an API provider built around predictable token costs, this is one of the stronger current options.
| Model Name | Best Use Case | Context Window | Input Price (per 1M tokens) | Output Price (per 1M tokens) |
| DeepSeek-V4-Pro | High-end reasoning, coding, and agent workflows | 1M | $1.30 | $2.60 |
Why this matters: On DeepInfra, DeepSeek V4 is priced at $1.30 per 1M input tokens and $2.60 per 1M output tokens. That gives you a very cost-efficient path for large-scale reasoning and coding workloads on a provider that also exposes cached tokens at $0.10 per 1M, which can further reduce spend when you reuse long prompts or repeated context.
If you expect heavy traffic, repeated-context prompts, or just want tighter control over serving costs, DeepInfra is one of the strongest places to run DeepSeek V4 in production. Teams comparing options across the broader model catalog often land on V4 Pro after weighing capability against price.
Below are practical developer scenarios where DeepInfra is a particularly strong way to run DeepSeek-V4-Pro: not because it is uniquely the cheapest in every benchmark, but because it combines the low benchmarked price tier with clear input/output pricing, cached-token pricing at $0.145 per 1M, JSON mode, function calling, and private endpoint support.
Scenario 1: RAG support bot with a large repeated knowledge prefix
If you are building a support copilot or internal docs assistant, you often resend the same long system prompt, retrieval scaffold, and tool instructions over and over. This is exactly the kind of workload where DeepInfra’s cached-token pricing is useful.
Assumptions
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 1,000 requests | DeepSeek-V4-Pro | DeepInfra | 20M standard input + 60M cached | 4M output | $42.40/month |
Cost breakdown
Why DeepInfra fits
Comparison: On a provider charging standard rates for all input at $1.30 per 1M with no cached-token discount, the same workload would cost $114.40/month, so DeepInfra saves $72/month.
Scenario 2: Code review and patch generation assistant
For a coding assistant that reads diffs, repository context, and issue text, then emits structured review comments or patch suggestions, DeepSeek-V4-Pro is attractive on capability alone. DeepInfra makes it easier to run that workload with predictable pricing and tool-friendly output.
Assumptions
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 10,000 requests | DeepSeek-V4-Pro | DeepInfra | 120M input | 20M output | $208/month |
Cost breakdown
Why DeepInfra fits
Comparison: On Azure at $1.91 per 1M input and $3.83 per 1M output, this 140M-token monthly workload would cost $305.80/month, so DeepInfra is $97.80/month cheaper.
Scenario 3: Agent workflow with persistent tool schemas and system instructions
Agent systems often pay an invisible tax: they keep resending the same policy text, tool definitions, and orchestration instructions. DeepInfra is a good fit when that repeated prompt overhead is real and not just theoretical. As DeepInfra’s role as a Hugging Face Inference Provider shows, chat completion and text generation tasks on open-weight LLMs like DeepSeek V4 are first-class workloads on the platform.
Assumptions
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 50,000 requests | DeepSeek-V4-Pro | DeepInfra | 200M standard input + 300M cached | 50M output | $420/month |
Cost breakdown
Why DeepInfra fits
Comparison: If all 500M input tokens were billed at the standard $1.30 per 1M input rate with no cache discount, the same workload would cost $780/month, so DeepInfra saves $360/month.
Scenario 4: Long-context document analysis pipeline
If you are processing contracts, research bundles, policy sets, or large case files, DeepSeek V4’s long-context design is a reason to care, but DeepInfra’s pricing structure is what makes repeated production use more manageable. For workloads that include scanned documents alongside text, you can also pair the V4 Pro pipeline with multimodal options like DeepSeek-OCR, which uses DeepEncoder and DeepSeek3B-MoE-A570M to extract structure before reasoning.
Assumptions
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 2,000 requests | DeepSeek-V4-Pro | DeepInfra | 200M input | 10M output | $286/month |
Cost breakdown
Why DeepInfra fits
Comparison: On Azure at $1.91 per 1M input and $3.83 per 1M output, this 210M-token workload would cost $420.30/month, so DeepInfra is $134.30/month cheaper.
Scenario 5: Private enterprise deployment for internal engineering tools
Sometimes the key advantage is not raw throughput. It is being able to use the same model in a more controlled deployment setup while keeping pricing understandable. That is where DeepInfra’s private endpoint support becomes more relevant than a leaderboard win. Teams that need dedicated compute can also spin up GPU instances and get from idea to a GPU-powered container in under 10 seconds, which is useful when you want isolation without long provisioning cycles.
Assumptions
| Volume | Model | Provider | Input Tokens | Output Tokens | Monthly Cost |
| 25,000 requests | DeepSeek-V4-Pro | DeepInfra | 200M input | 37.5M output | $357.50/month |
Cost breakdown
Why DeepInfra fits
Comparison: On Azure at $1.91 per 1M input and $3.83 per 1M output, this 237.5M-token workload would cost $525.63/month, so DeepInfra is $168.13/month cheaper.
| Scenario | Monthly Requests | Monthly Cost on DeepInfra | Primary Cost Driver |
| RAG support bot | 1,000 | $42.40 | Cached input discount |
| Code review assistant | 10,000 | $208.00 | Standard input and output volume |
| Agent workflow | 50,000 | $420.00 | Cached input discount at scale |
| Document analysis | 2,000 | $286.00 | High input volume per request |
| Private enterprise deployment | 25,000 | $357.50 | Output volume and private endpoint needs |
Choosing a provider for DeepSeek V4 Pro is less about finding the “best” option and more about matching provider economics to how your workload actually behaves. The model itself is strong across reasoning, coding, and long-context tasks regardless of where you run it. What differs meaningfully between providers is how you pay for that capability — and whether the platform gives you the controls to keep costs predictable as usage grows.
The two criteria that matter most in practice are token pricing structure and prompt caching. If your app resends large system prompts, tool definitions, or retrieval context repeatedly, the difference between a provider that discounts cached tokens and one that does not is real money, not a theoretical saving. DeepInfra’s cached rate of $0.10 per 1M tokens is among the lowest currently available for this model. Beyond caching, output token cost deserves more attention than it usually gets. DeepSeek V4 Pro is the kind of model people use for long code generation and multi-step reasoning, and those workloads generate output tokens fast. At $2.60 per 1M output tokens, verbosity still adds up, and providers that support JSON mode and function calling help you avoid the retry loops that quietly inflate output counts.
DeepInfra also covers the deployment concerns that matter once you move past early testing: SOC 2 and ISO 27001 certification, private endpoint support, and bare-metal infrastructure that removes a layer of overhead. You can review the V4 Pro API reference to see how straightforward integration looks, or browse the full text generation model catalog if you want to compare DeepSeek V4 Pro against other options before committing. The pricing is transparent, the infrastructure is production-ready, and the first call is easy to make.
DeepSeek V4 Pro costs $1.30 per 1M input tokens, $2.60 per 1M output tokens, and $0.10 per 1M cached tokens on DeepInfra’s standard tier, as of October 2026.
It depends on which provider OpenRouter routes your request to. OpenRouter lists per-provider pricing for this model, and DeepInfra’s rate is among the lower options in that list, but OpenRouter can also route to providers priced above or below DeepInfra.
Yes. DeepInfra bills cached input at $0.10 per 1M tokens, a 92% discount against its standard input rate, and most other tracked providers now publish a cache price as well.
DeepSeek V4 Pro supports a 1,048,576-token context window, commonly listed as 1M tokens.
Yes. DeepInfra lists DeepSeek V4 Pro under the MIT license, and the model weights are available on Hugging Face.
DeepInfra’s model page lists DeepSeek V4 Pro as zero retention, meaning request data is not stored after the response is returned.
Introducing the Flex Service Tier: Cheaper Inference When You Can WaitRun latency-tolerant work at 0.8× real-time — best-effort, sheddable, same OpenAI-compatible API.
DeepInfra Launches Access to NVIDIA Nemotron Models for Vision, Retrieval, and AI SafetyDeepInfra is serving the new, open NVIDIA Nemotron vision language and OCR AI models from day zero of their release. As a leading inference provider committed to performance and cost-efficiency, we're making these cutting-edge models available at the industry's best prices, empowering developers to build specialized AI agents without compromising on budget or performance.
Beat AI Subscription Fatigue With One API<p>Open your company card statement and scroll the recurring charges. Twenty dollars for a chat assistant, twenty more for a coding copilot, fifteen for an image API, another forty for the automation glue that wires them together. None of them is expensive on its own. Together they are a slow leak you stopped noticing months […]</p>
© 2026 DeepInfra. All rights reserved.