DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

DeepSeek V4 Pro Pricing Guide 2026: Pricing, Providers & Cost Comparison

Published on 2026.04.30 by DeepInfra
DeepSeek V4 Pro Pricing Guide 2026: Pricing, Providers & Cost Comparison

DeepSeek V4 Pro is an open-weight Mixture-of-Experts model with 1.6T total parameters, 49B active parameters, and a 1M-token context window. It ships under the MIT license and supports JSON mode and function calling. The V4 Pro 0813 release is the GA version of the model first previewed on April 24, 2026.

DeepSeek V4 Pro costs $1.30 per 1M input tokens, $2.60 per 1M output tokens, and $0.10 per 1M cached tokens on DeepInfra. Those rates sit 25% below the previous DeepInfra pricing of $1.74 input and $3.48 output, and cached input is 31% cheaper than the earlier $0.145.

Artificial Analysis now tracks 10 API providers for the model. DeepInfra ranks second on blended price at $0.59 per 1M tokens, behind Systalyze at $0.33. Output speed ranges from 56 to 279.9 tokens per second across providers, so cost and latency requirements should drive your choice. This guide covers current pricing, provider benchmarks, service tiers, and five worked monthly cost scenarios.

For developers and ML teams, the practical question is not whether DeepSeek V4 Pro is interesting. It is whether you want the cheapest access path, the fastest output speed, the lowest raw time to first token, or a platform setup that fits your stack. The benchmark data here makes that evaluation unusually concrete: Fireworks leads on output speed and end-to-end latency, DeepInfra ties for the lowest benchmarked price while adding cached-token pricing and private deployment support, Together.ai posts the best raw time to first token at 0.99 seconds, and OpenRouter offers another access path with much lower listed per-token rates on its model page.

DeepSeek V4 Executive Summary

DeepSeek V4 Pro is an open-weight reasoning model for coding, long-context analysis, and agent workflows. Artificial Analysis benchmarks it across 10 API providers, and DeepInfra ranks second on blended price at $0.59 per 1M tokens. DeepInfra lists $1.30 input, $2.60 output, and $0.10 cached per 1M tokens, with optional Priority and Flex tiers. Choose DeepInfra when cost control, caching, and data retention matter most. Choose a faster provider when response time is the main constraint. 

Best ForProvider RecommendationWhy
Cost-sensitive production workloadsDeepInfraSecond-lowest blended price on Artificial Analysis at $0.59 per 1M tokens, with list rates of $1.30 input and $2.60 output.
Single API across many providersOpenRouterLists 16 providers for the model and can reroute requests to another provider when one returns errors.
Lowest first-chunk latencyBasetenMedian first chunk arrives in 0.65 seconds on Artificial Analysis, ahead of DigitalOcean at 0.84 and DeepInfra at 1.23.
RAG and repeated-context workloadsDeepInfraCached input costs $0.10 per 1M tokens, a 92% discount to the $1.30 standard input rate.
Maximum output speed and lowest answer latencySystalyzeGenerates 279.9 tokens per second and returns the first answer token in 8.65 seconds, ahead of Crusoe and DeepSeek.
Lowest blended priceSystalyzePosts a $0.33 blended price per 1M tokens on a 7:2:1 cache, input, and output mix, ahead of DeepInfra at $0.59.
Batch and non-urgent workloadsDeepInfra FlexThe Flex tier bills $1.04 input and $2.08 output per 1M tokens, a 20% discount, in exchange for best-effort scheduling.
Direct access from the model creatorDeepSeekThe first-party API lists $1.32 input and $3.96 output per 1M tokens and runs at 109.9 tokens per second.
Latency-sensitive workloads on DeepInfraDeepInfra PriorityThe Priority tier bills $1.95 input and $3.90 output per 1M tokens for faster first tokens and higher throughput at peak demand.
Data retention and compliance controlsDeepInfraZero retention, SOC 2 and ISO 27001 certification, and private endpoint deployment.

Sources: Artificial Analysis (DeepSeek V4 Pro 0813, Max reasoning effort, 10,000-token workload), plus the DeepInfra and OpenRouter model pages, accessed October 2026.

What Changed in DeepSeek V4 Pro Pricing and Availability

Several details of DeepSeek V4 Pro pricing and availability have shifted since its launch. The table below summarizes the current picture.

AreaEarlier version of this guideCurrent
DeepInfra input price$1.74 per 1M tokens$1.30 per 1M tokens
DeepInfra output price$3.48 per 1M tokens$2.60 per 1M tokens
DeepInfra cached input price$0.145 per 1M tokens$0.10 per 1M tokens
Serving precisionFP4fp8, per the DeepInfra model page
Model releaseApril 24, 2026 previewV4 Pro 0813 GA, released August 13, 2026
DeepInfra service tiersStandard onlyStandard, Priority (1.5×), and Flex (0.8×)
Providers benchmarked610
Blended price method3:1 input-to-output ratio7:2:1 cache, input, and output ratio
DeepInfra blended price$2.17 per 1M tokens$0.59 per 1M tokens
Cached-token pricingListed for DeepInfra onlyListed for most providers on OpenRouter and Artificial Analysis

The two blended prices use different ratios, so they are not directly comparable. The current Artificial Analysis blend weights cached input at 70%, which is why blended figures now sit well below list prices.

Understanding Tokens and How You’re Charged

If you have ever looked at a model bill and thought, “that prompt was not that long,” this is the part that explains why it was.

A token is a chunk of text the model reads or generates. It is not the same thing as a word. Short words may be one token. Longer words, code, JSON, whitespace patterns, and weird punctuation splits can turn into more. Pricing is based on tokens, so the shape of your workload matters more than the raw character count.

  • Input tokens are everything you send to the model.
  • Your system prompt
  • User message
  • Tool schemas
  • Retrieved context
  • Chat history
  • Any hidden scaffolding your app injects
  • Output tokens are everything the model sends back.
  • Natural language answers
  • Code
  • JSON objects
  • Function-call arguments
  • In reasoning models, potentially a lot of generated text if you let responses run long
  • Cached tokens apply when a provider discounts repeated prompt content.
  • This usually helps with large static prefixes
  • Think long system prompts, repeated documents, or persistent agent instructions
  • Not every provider exposes this as a separate line item
  • Blended price is the shortcut number many benchmarks use.
  • Artificial Analysis now compares providers using a 7:2:1 cache hit, input, and output ratio
  • This ratio assumes a cache-heavy workload, which favors providers with strong cache discounts
  • It is not your real bill unless your workload actually matches that ratio
Token typeWhat it isWhy it matters
Input tokensTokens you send in the request promptUsually the biggest driver for RAG, long-context, and agent workloads
Output tokensTokens the model generates backExpensive when you ask for long answers, code diffs, or verbose JSON
Cached tokensReused prompt tokens billed at a reduced rate when supportedCan materially cut cost for repeated context-heavy requests
Blended tokensA benchmarked combination of cache, input, and output token costsGood for provider comparison, bad for forecasting if your traffic mix is different

Provider Token Cost Tradeoffs for DeepSeek V4

DeepSeek V4 is one of those models where the “headline price” can mislead you if you do not check how the provider bills input versus output.

ProviderToken cost profileAdvantageDisadvantage
DeepInfra$1.30 input / $2.60 output / $0.10 cached per 1MSecond-lowest blended price in current provider data, with the largest cache discount at 92% off standard inputOutput speed trails faster providers at under 60 tokens per second
DeepSeek (first-party)$1.32 input / $3.96 output / $0.044 cached per 1MAccess directly from the model creator, with output speed over 100 tokens per secondOutput price runs higher than DeepInfra and most third-party providers
Parasail$0.45 input / $3.48 output / $0.10 cached per 1MLowest input price among currently tracked providers, with sub-second latencyOutput price sits close to DeepInfra despite the lower input rate
DigitalOcean$1.04 input / $2.09 output / $0.21 cached per 1MHighest reported uptime among tracked providers at 99.91%Cached input price runs about double DeepInfra’s rate
SiliconFlow$1.50 input / $3.14 output / $0.14 cached per 1MSpecs mirror the official DeepSeek API, useful as a routing fallbackHigher latency than DeepInfra at just over 2 seconds to first token
Novita$1.60 input / $3.20 output / $0.14 cached per 1MComparable throughput to DeepInfra, with full 1M contextInput and output prices both run higher than DeepInfra
Azure$1.91 input / $3.83 output / $0.16 cached per 1MEnterprise billing and compliance tooling already in place for Azure customersHighest token pricing among the providers compared here
OpenRouterPricing varies by routed provider, from about $0.21 to $1.91 input per 1MAutomatic failover across providers if one returns an errorCache and cost depend on which underlying provider handles the request

Source: OpenRouter provider listing for deepseek/deepseek-v4-pro and Artificial Analysis, accessed October 2026.

DeepInfra and Parasail show the most useful cache discounts in the current provider data.

  • DeepInfra cuts cached input to $0.10 per 1M tokens, a 92% discount against its own $1.30 standard rate
  • Parasail lists the same $0.10 cached rate, paired with the lowest input price in the comparison at $0.45 per 1M tokens
  • Apps that resend long system prompts or retrieval context benefit most from either option

Base token rates now differ meaningfully across providers, unlike the earlier benchmark set.

  • DeepInfra’s $1.30 input and $2.60 output per 1M tokens sit below DigitalOcean, SiliconFlow, Novita, and Azure on both figures
  • Parasail undercuts DeepInfra on input price but charges a similar output rate
  • Provider choice now affects your bill more directly than it did when several providers clustered at the same price point

Output tokens are where teams get sloppy and bills get weird.

  • DeepSeek V4 Pro is a strong coding and reasoning model
  • Strong reasoning models tend to be used for long answers, large code generations, structured outputs, and multi-step agent traces
  • Output prices across the compared providers range from $2.09 to $3.96 per 1M tokens, so a verbose app can shift your monthly bill by close to double

OpenRouter pricing now tracks each underlying provider directly.

  • The OpenRouter model page lists per-provider input, output, and cache pricing rather than one effective rate
  • Prices on OpenRouter range from roughly $0.21 to $1.91 per 1M input tokens depending on which provider handles the request
  • Routing mode (Balanced, Nitro, Floor, or Exacto) changes which provider you land on, so check the routing setting before assuming a listed price applies

Cache-token pricing is no longer a DeepInfra-only feature.

  • Most providers tracked on OpenRouter and Artificial Analysis now list a cache read price alongside input and output
  • DeepInfra and Parasail currently offer the lowest cached rate at $0.10 per 1M tokens among the providers compared here
  • Confirm whether a provider bills cache writes separately, since several only publish the cache read price

Practical rule of thumb for DeepSeek V4 Pro billing

  • If you send large repeated prompts: cached-token pricing matters most, and DeepInfra or Parasail currently lead on that rate
  • If you generate long answers or code: output-token pricing matters most, and the gap between providers now runs close to double
  • If your prompt and response mix is ordinary: compare blended prices directly, since providers no longer cluster at the same rate
  • If you use OpenRouter: check which provider is serving your request, since pricing and cache support vary by provider rather than by a single platform-wide rate

Practical rule of thumb for DeepSeek V4 billing

  • If you send huge repeated prompts: cached-token pricing matters most
  • If you generate long answers or code: output-token pricing matters most
  • If your prompt/response mix is ordinary: the benchmarked low-cost providers are close enough that latency will probably decide it
  • If you use OpenRouter: verify actual routed cost under your traffic pattern instead of trusting the listing blindly

Reasoning Effort and Output Token Cost

DeepSeek V4 Pro supports three reasoning modes: Non-think, Think High, and Think Max. Each trades response quality against the number of tokens the model generates before answering.

Non-think mode returns fast, direct responses with no visible reasoning step, suited to routine, low-risk tasks. Think High mode adds a visible reasoning pass for complex problem-solving and planning. Think Max pushes reasoning effort to its limit for the hardest tasks, at the cost of the most output tokens.

Reasoning tokens bill at the same output rate as the final answer. On DeepInfra, that means $2.60 per 1M tokens on standard pricing, so a Think Max request can cost meaningfully more than the same prompt run in Non-think mode. Artificial Analysis reports that DeepSeek V4 Pro 0813 at max reasoning effort generated 160 million output tokens across its Intelligence Index evaluation, above the median for comparable open-weight models.

If your app defaults to a high reasoning effort for every request, confirm that the task actually needs it. Routine classification, extraction, or short-answer tasks rarely benefit from Think Max and will cost more than necessary if it runs by default.

Choosing Between DeepSeek V4 Pro, V4 Flash, and V4.1 Flash

DeepSeek V4 Pro is the largest model in the family, built for workloads that need the strongest reasoning and coding results regardless of cost. DeepSeek V4 Flash trades some capability for lower cost and faster inference. DeepSeek V4.1 Flash is the newer multimodal release, built on a different architecture that activates far fewer parameters per token and adds native image understanding.

If your workload is reasoning-heavy, long-context, or code-focused and cost is a secondary concern, V4 Pro remains the stronger pick. If you need lower cost per token, faster throughput, or image input, V4.1 Flash is worth evaluating directly against V4 Pro for your specific task before committing. See the [DeepSeek V4.1 Flash pricing guide] for a full cost breakdown of that model.

DeepInfra: the power user’s choice for DeepSeek V4

DeepInfra is a strong choice for DeepSeek V4 Pro because it combines straightforward token pricing with infrastructure built for production deployment. The platform runs on bare-metal infrastructure, which removes a layer of virtualization overhead that can affect both performance consistency and cost. Standard pricing sits at $1.30 per 1M input tokens, $2.60 per 1M output tokens, and $0.10 per 1M cached tokens, with Priority and Flex tiers available for workloads that need faster or cheaper service. If you want an API provider built around predictable token costs, this is one of the stronger current options.

Model NameBest Use CaseContext WindowInput Price (per 1M tokens)Output Price (per 1M tokens)
DeepSeek-V4-ProHigh-end reasoning, coding, and agent workflows1M$1.30$2.60

Why this matters: On DeepInfra, DeepSeek V4 is priced at $1.30 per 1M input tokens and $2.60 per 1M output tokens. That gives you a very cost-efficient path for large-scale reasoning and coding workloads on a provider that also exposes cached tokens at $0.10 per 1M, which can further reduce spend when you reuse long prompts or repeated context.

If you expect heavy traffic, repeated-context prompts, or just want tighter control over serving costs, DeepInfra is one of the strongest places to run DeepSeek V4 in production. Teams comparing options across the broader model catalog often land on V4 Pro after weighing capability against price.

Real-world cost scenarios for developers

Below are practical developer scenarios where DeepInfra is a particularly strong way to run DeepSeek-V4-Pro: not because it is uniquely the cheapest in every benchmark, but because it combines the low benchmarked price tier with clear input/output pricing, cached-token pricing at $0.145 per 1M, JSON mode, function calling, and private endpoint support.

Scenario 1: RAG support bot with a large repeated knowledge prefix

If you are building a support copilot or internal docs assistant, you often resend the same long system prompt, retrieval scaffold, and tool instructions over and over. This is exactly the kind of workload where DeepInfra’s cached-token pricing is useful.

Assumptions

  • Volume: 1,000 requests/month
  • Model: DeepSeek-V4-Pro
  • Provider: DeepInfra
  • Per request: 80,000 input tokens, 4,000 output tokens
  • Of the input, 60,000 tokens are cached/reused; 20,000 tokens billed at the normal input rate
VolumeModelProviderInput TokensOutput TokensMonthly Cost
1,000 requestsDeepSeek-V4-ProDeepInfra20M standard input + 60M cached4M output$42.40/month

Cost breakdown

  • Standard input: 20M × $1.30/1M = $26
  • Cached input: 60M × $0.10/1M = $6
  • Output: 4M × $2.60/1M = $10.40
  • Total: $42.40/month

Why DeepInfra fits

  • Reused long prompt context is common in RAG.
  • Cached tokens at $0.145 per 1M are the standout lever here.
  • You still get JSON mode and function calling for structured retrieval and tool use.

Comparison: On a provider charging standard rates for all input at $1.30 per 1M with no cached-token discount, the same workload would cost $114.40/month, so DeepInfra saves $72/month.

Scenario 2: Code review and patch generation assistant

For a coding assistant that reads diffs, repository context, and issue text, then emits structured review comments or patch suggestions, DeepSeek-V4-Pro is attractive on capability alone. DeepInfra makes it easier to run that workload with predictable pricing and tool-friendly output.

Assumptions

  • Volume: 10,000 requests/month
  • Model: DeepSeek-V4-Pro
  • Provider: DeepInfra
  • Per request: 12,000 input tokens, 2,000 output tokens
VolumeModelProviderInput TokensOutput TokensMonthly Cost
10,000 requestsDeepSeek-V4-ProDeepInfra120M input20M output$208/month

Cost breakdown

  • Input: 120M × $1.30/1M = $156.00
  • Output: 20M × $2.60/1M = $52.00
  • Total: $208.00/month

Why DeepInfra fits

  • DeepSeek-V4-Pro posts 93.5 on LiveCodeBench, 3206 on Codeforces, and 80.6 on SWE Verified, which is exactly the profile many code-assistant builders want.
  • JSON mode is useful for returning review findings in a stable schema.
  • Function calling helps when the assistant needs to trigger CI checks, repo lookups, or patch workflows.
  • If you later need stronger isolation, private endpoint deployment is available.

Comparison: On Azure at $1.91 per 1M input and $3.83 per 1M output, this 140M-token monthly workload would cost $305.80/month, so DeepInfra is $97.80/month cheaper.

Scenario 3: Agent workflow with persistent tool schemas and system instructions

Agent systems often pay an invisible tax: they keep resending the same policy text, tool definitions, and orchestration instructions. DeepInfra is a good fit when that repeated prompt overhead is real and not just theoretical. As DeepInfra’s role as a Hugging Face Inference Provider shows, chat completion and text generation tasks on open-weight LLMs like DeepSeek V4 are first-class workloads on the platform.

Assumptions

  • Volume: 50,000 requests/month
  • Model: DeepSeek-V4-Pro
  • Provider: DeepInfra
  • Per request: 10,000 total input tokens (6,000 cached, 4,000 standard), 1,000 output tokens
VolumeModelProviderInput TokensOutput TokensMonthly Cost
50,000 requestsDeepSeek-V4-ProDeepInfra200M standard input + 300M cached50M output$420/month

Cost breakdown

  • Standard input: 200M × $1.30/1M = $260.00
  • Cached input: 300M × $0.10/1M = $30.00
  • Output: 50M × $2.60/1M = $130.00
  • Total: $420.00/month

Why DeepInfra fits

  • This is the classic “agent loop with repeated scaffolding” pattern.
  • Cached-token billing lowers the cost of persistent instructions and tool schemas.
  • Function calling is table stakes for agent execution, and DeepInfra supports it.
  • DeepSeek-V4-Pro also scores 73.6 on MCPAtlas Public and 51.8 on Toolathlon, which makes the model itself relevant for tool-using systems.

Comparison: If all 500M input tokens were billed at the standard $1.30 per 1M input rate with no cache discount, the same workload would cost $780/month, so DeepInfra saves $360/month.

Scenario 4: Long-context document analysis pipeline

If you are processing contracts, research bundles, policy sets, or large case files, DeepSeek V4’s long-context design is a reason to care, but DeepInfra’s pricing structure is what makes repeated production use more manageable. For workloads that include scanned documents alongside text, you can also pair the V4 Pro pipeline with multimodal options like DeepSeek-OCR, which uses DeepEncoder and DeepSeek3B-MoE-A570M to extract structure before reasoning.

Assumptions

  • Volume: 2,000 requests/month
  • Model: DeepSeek-V4-Pro
  • Provider: DeepInfra
  • Per request: 100,000 input tokens, 5,000 output tokens
VolumeModelProviderInput TokensOutput TokensMonthly Cost
2,000 requestsDeepSeek-V4-ProDeepInfra200M input10M output$286/month

Cost breakdown

  • Input: 200M × $1.30/1M = $260.00
  • Output: 10M × $2.60/1M = $26.00
  • Total: $286.00/month

Why DeepInfra fits

  • DeepSeek-V4-Pro-Max scores 83.5 on MRCR 1M and 62.0 on CorpusQA 1M, so the model is built for serious long-context use.
  • DeepInfra exposes clear per-token pricing instead of only a blended benchmark number.
  • If your pipeline has repeated boilerplate instructions or stable analysis templates, cached tokens can matter in later optimization passes.
  • For document-heavy pipelines, the broader multimodal model lineup gives you OCR and vision options to feed structured text into V4 Pro.

Comparison: On Azure at $1.91 per 1M input and $3.83 per 1M output, this 210M-token workload would cost $420.30/month, so DeepInfra is $134.30/month cheaper.

Scenario 5: Private enterprise deployment for internal engineering tools

Sometimes the key advantage is not raw throughput. It is being able to use the same model in a more controlled deployment setup while keeping pricing understandable. That is where DeepInfra’s private endpoint support becomes more relevant than a leaderboard win. Teams that need dedicated compute can also spin up GPU instances and get from idea to a GPU-powered container in under 10 seconds, which is useful when you want isolation without long provisioning cycles.

Assumptions

  • Volume: 25,000 requests/month
  • Model: DeepSeek-V4-Pro
  • Provider: DeepInfra
  • Per request: 8,000 input tokens, 1,500 output tokens
VolumeModelProviderInput TokensOutput TokensMonthly Cost
25,000 requestsDeepSeek-V4-ProDeepInfra200M input37.5M output$357.50/month

Cost breakdown

  • Input: 200M × $1.30/1M = $260.00
  • Output: 37.5M × $2.60/1M = $97.50
  • Total: $357.50/month

Why DeepInfra fits

  • Good match for internal copilots, secure workflow assistants, or engineering automation where a private endpoint matters.
  • DeepInfra is SOC 2 Certified and ISO 27001 Certified.
  • You still retain JSON mode and function calling for production integrations.
  • If your assistant also needs voice output, you can extend the stack with options from the text-to-speech model catalog without leaving the same provider.

Comparison: On Azure at $1.91 per 1M input and $3.83 per 1M output, this 237.5M-token workload would cost $525.63/month, so DeepInfra is $168.13/month cheaper.

Monthly Cost Summary Across Scenarios

ScenarioMonthly RequestsMonthly Cost on DeepInfraPrimary Cost Driver
RAG support bot1,000$42.40Cached input discount
Code review assistant10,000$208.00Standard input and output volume
Agent workflow50,000$420.00Cached input discount at scale
Document analysis2,000$286.00High input volume per request
Private enterprise deployment25,000$357.50Output volume and private endpoint needs

Conclusion

Choosing a provider for DeepSeek V4 Pro is less about finding the “best” option and more about matching provider economics to how your workload actually behaves. The model itself is strong across reasoning, coding, and long-context tasks regardless of where you run it. What differs meaningfully between providers is how you pay for that capability — and whether the platform gives you the controls to keep costs predictable as usage grows.

The two criteria that matter most in practice are token pricing structure and prompt caching. If your app resends large system prompts, tool definitions, or retrieval context repeatedly, the difference between a provider that discounts cached tokens and one that does not is real money, not a theoretical saving. DeepInfra’s cached rate of $0.10 per 1M tokens is among the lowest currently available for this model. Beyond caching, output token cost deserves more attention than it usually gets. DeepSeek V4 Pro is the kind of model people use for long code generation and multi-step reasoning, and those workloads generate output tokens fast. At $2.60 per 1M output tokens, verbosity still adds up, and providers that support JSON mode and function calling help you avoid the retry loops that quietly inflate output counts.

DeepInfra also covers the deployment concerns that matter once you move past early testing: SOC 2 and ISO 27001 certification, private endpoint support, and bare-metal infrastructure that removes a layer of overhead. You can review the V4 Pro API reference to see how straightforward integration looks, or browse the full text generation model catalog if you want to compare DeepSeek V4 Pro against other options before committing. The pricing is transparent, the infrastructure is production-ready, and the first call is easy to make.

Frequently Asked Questions

How much does DeepSeek V4 Pro cost on DeepInfra?

DeepSeek V4 Pro costs $1.30 per 1M input tokens, $2.60 per 1M output tokens, and $0.10 per 1M cached tokens on DeepInfra’s standard tier, as of October 2026.

Is DeepInfra cheaper than OpenRouter for DeepSeek V4 Pro?

It depends on which provider OpenRouter routes your request to. OpenRouter lists per-provider pricing for this model, and DeepInfra’s rate is among the lower options in that list, but OpenRouter can also route to providers priced above or below DeepInfra.

Does DeepSeek V4 Pro support prompt caching?

Yes. DeepInfra bills cached input at $0.10 per 1M tokens, a 92% discount against its standard input rate, and most other tracked providers now publish a cache price as well.

What is the context window for DeepSeek V4 Pro?

DeepSeek V4 Pro supports a 1,048,576-token context window, commonly listed as 1M tokens.

Is DeepSeek V4 Pro open source?

Yes. DeepInfra lists DeepSeek V4 Pro under the MIT license, and the model weights are available on Hugging Face.

Does DeepInfra retain my data when I use DeepSeek V4 Pro?

DeepInfra’s model page lists DeepSeek V4 Pro as zero retention, meaning request data is not stored after the response is returned.

Related articles
Introducing the Flex Service Tier: Cheaper Inference When You Can WaitIntroducing the Flex Service Tier: Cheaper Inference When You Can WaitRun latency-tolerant work at 0.8× real-time — best-effort, sheddable, same OpenAI-compatible API.
DeepInfra Launches Access to NVIDIA Nemotron Models for Vision, Retrieval, and AI SafetyDeepInfra Launches Access to NVIDIA Nemotron Models for Vision, Retrieval, and AI SafetyDeepInfra is serving the new, open NVIDIA Nemotron vision language and OCR AI models from day zero of their release. As a leading inference provider committed to performance and cost-efficiency, we're making these cutting-edge models available at the industry's best prices, empowering developers to build specialized AI agents without compromising on budget or performance.
Beat AI Subscription Fatigue With One APIBeat AI Subscription Fatigue With One API<p>Open your company card statement and scroll the recurring charges. Twenty dollars for a chat assistant, twenty more for a coding copilot, fifteen for an image API, another forty for the automation glue that wires them together. None of them is expensive on its own. Together they are a slow leak you stopped noticing months [&hellip;]</p>