DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Qwen3.8-27B Pricing: DeepInfra vs Alibaba API
Published on 2026.10.06 by DeepInfra
Qwen3.8-27B Pricing: DeepInfra vs Alibaba API

Qwen3.8-27B is Alibaba’s 27B open-weight reasoning model, released August 14, 2026 under Apache 2.0. It ships with a large context window, multimodal input, and API availability across eight providers. Pricing varies sharply by host: Alibaba’s first-party rate runs over 3x DeepInfra’s for the same weights.

Context window figures differ slightly by source. Artificial Analysis reports 256K tokens; DeepInfra lists a native 262,144, extensible to 1,000,000. Both describe the same model: text, image, and video input, text-only output, and public weights on Hugging Face, with JSON output, function calling, and multimodal workflows supported.

Artificial Analysis gives it an Intelligence Index score of 34, against a median of 8 for comparable open-weight models, well above average in its class. DeepInfra’s benchmark page shows the same strength against Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B, and Opus4.6 Max. It scores 73.0 on Terminal Bench 2.1, 61.7 on SWE-bench Pro, 84.3 on OSWorld-Verified, and 64.8 on WebArena-Verified. The tradeoff is efficiency. Artificial Analysis classifies the model as notably slow and very verbose: 44.8 output tokens per second, 3.88 seconds to first token, and 200M output tokens generated during Intelligence Index evaluation.

This verbosity makes provider choice matter more than usual. Alibaba’s API prices Qwen3.8-27B at $0.50 per 1M input tokens and $3.00 per 1M output tokens, rates Artificial Analysis calls expensive relative to similar open-weight models. DeepInfra lists the same model at $0.15 input, $1.875 output, and $0.038 cached per 1M, a meaningfully lower bill for long-context RAG, multimodal agents, and sustained production traffic.

Qwen3.8-27B Executive Summary

Qwen3.8-27B looks strongest when you want a high-performing open-weight reasoning model and care about deployment flexibility as much as raw benchmarks. The pricing spread is wide. Alibaba’s API is listed at $0.50 input / $3.00 output per 1M tokens, while DeepInfra lists $0.15 input / $1.875 output / $0.038 cached. The best provider choice depends less on the model itself than on whether you prioritize first-party access or lower-cost production inference. Compared with Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B, and Opus4.6 Max, Qwen3.8-27B is a serious contender for coding, agent, and multimodal workloads, but verbosity and latency are worth watching.

Best ForProviderWhy
Lowest price / cost-sensitive workloadsDeepInfraDeepInfra’s standard pricing of $0.15 input and $1.875 output per 1M tokens is lower than Alibaba’s $0.50 input and $3.00 output for the same model.
RAG, document-heavy, or high-throughput use casesDeepInfraDeepInfra combines a 262,144-token public context window with cached-token pricing at $0.038 per 1M, which fits long-context and repeat-query workloads.
Easiest onboarding / fastest time-to-first-callDeepInfraThe model is deployment-ready on DeepInfra’s inference cloud with public endpoint access, JSON output, function calling, multimodal input, and private endpoint availability.
Multimodal agents and tool-using applicationsDeepInfraDeepInfra exposes image and video input plus function calling, with strong results on OSWorld-Verified, WebArena-Verified, AndroidWorld, and Vision2Web.
First-party managed accessAlibaba APIArtificial Analysis sources its pricing and latency metrics directly from Alibaba’s first-party API, the clearest reference point for official hosted access.
Benchmark-first evaluation against nearby alternativesDeepInfraDeepInfra provides side-by-side benchmark comparisons versus Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B, and Opus4.6 Max across coding, agent, general, and multimodal tasks.

Understanding Tokens and How You’re Charged

Tokens are the billing unit for LLM APIs, and they are not the same thing as words. A token can be a whole short word, part of a long word, punctuation, JSON syntax, a code symbol, or a whitespace pattern. For practical budgeting: 1M tokens is a lot of text. Production bills usually come from repetition, long prompts, agent loops, and verbose outputs, more than from a single big request.

Token TypeWhat It IsWhy It Matters
Input tokensEverything you send to the model: system prompt, user prompt, chat history, tool schemas, retrieved docs, and images or video encoded into the request.This is the part that quietly grows over time. RAG, long chat threads, and large tool definitions can make input cost dominate faster than expected.
Output tokensEverything the model generates back: answers, JSON, code, tool-call arguments, and long-form completions.Qwen3.8-27B is very verbose in evaluation. If your app allows long answers, output cost can become the main line item.
Cached tokensReused prompt tokens the provider recognizes from previous requests and bills at a reduced rate.This matters most for stable prefixes: system prompts, repeated instructions, and recurring RAG scaffolding. Cache pricing can materially change total cost.

Token accounting matters more than usual for Qwen3.8-27B because the model combines a very large context window, multimodal input, reasoning support, and a tendency toward long outputs. It is a useful mix for agents and document-heavy workflows. It is also how a bill that looked fine in a spreadsheet looks less fine in production. A few cost behaviors worth keeping in mind:

  • Long context is not free just because you have it. A 256K to 262K context window is capacity, not a discount.
  • Verbose models shift spend toward output tokens. Artificial Analysis reports 200M output tokens generated during its evaluation run, against an 82M median for comparable models.
  • Multimodal requests can increase input usage. Image and video support is useful, but it still becomes billable input.
  • Tool use can inflate both sides. Function schemas add input tokens; structured outputs and tool arguments add output tokens.
  • Chat history drift is real. Appending prior messages increases token usage even when the user asks short questions.

Provider Token Cost Differences for Qwen3.8-27B

The headline difference is simple: the same model is much cheaper on DeepInfra than on Alibaba’s first-party API. The less obvious difference is where the savings show up: DeepInfra is lower on input tokens, lower on output tokens, and exposes a clearly listed cached-token rate. If you run long-context RAG, repeat a large system prompt, or keep agent instructions stable across calls, cache pricing matters more than most teams assume at the start.

ProviderInput TokensOutput TokensCached TokensNotes
DeepInfra$0.15 / 1M$1.875 / 1M$0.038 / 1MLowest listed cost in this comparison. Best fit for prompt-heavy apps, long-context RAG, and repeated workflows where cache hits are common.
Alibaba API$0.50 / 1M$3.00 / 1M90% cache discount off listClear first-party reference pricing. Much more expensive on both input and output, and output pricing is especially costly for a verbose model.

The provider choice changes the economics quickly:

  • Input cost gap: Alibaba is over 3x DeepInfra’s listed input price ($0.50 vs. $0.15).
  • Output cost gap: Alibaba is 60% higher than DeepInfra’s listed output price ($3.00 vs. $1.875).
  • Cache-sensitive workloads: DeepInfra’s explicit cached rate of $0.038 per 1M is attractive for repeated prefixes. Alibaba’s 90% cache discount helps too, but the uncached baseline is much higher to begin with.

How Different Workloads Feel This

  • Short prompt, short answer apps: The provider gap matters, but less dramatically. There are not enough tokens in flight for cache strategy to do much.
  • RAG and document analysis: DeepInfra has the cleaner pricing story. Large retrieved context hits input billing every call unless you deliberately reuse cached prefixes.
  • Agents with long instructions and tool schemas: Cache support matters a lot. Repeated scaffolding can become a major hidden cost if billed as full-price input every turn.
  • Verbose coding or reasoning workflows: Output pricing matters most. Qwen3.8-27B’s tendency to produce long completions makes Alibaba’s $3.00 output rate harder to justify unless you specifically need first-party hosting.

The practical read: if you care about token efficiency, DeepInfra is the easier default for Qwen3.8-27B. If you want first-party managed access, Alibaba is the reference option, but you pay a real premium for the same model’s tokens. If your app has high cache reuse, test both with real prompts, since cache behavior can change total cost more than raw input price suggests. If your app lets the model think out loud or produce long JSON and code responses, cap output length early, or Qwen3.8-27B will spend your budget for you.

DeepInfra: The Power User’s Choice for Qwen3.8-27B

DeepInfra runs inference on bare-metal infrastructure, which matters because cutting out extra virtualization layers helps with both performance consistency and cost efficiency. It is typically 50 to 80% cheaper than major cloud competitors, with appeal among developers, high-volume API users, and teams watching unit economics closely. For teams that want strong open-model capability without premium-hosting prices, browsing the full model catalog is a practical place to start.

DeepInfra lists Qwen3.8-27B at $0.15 per 1M input tokens and $1.875 per 1M output tokens, against $0.50 input and $3.00 output on Alibaba’s API for the same model, a meaningful reduction on both sides of the bill, especially for prompt-heavy workloads or long responses. If you expect sustained traffic, the provider choice directly changes whether this model feels affordable in production.

Real-World Cost Scenarios for Developers

Below are practical scenarios where DeepInfra is a particularly strong fit for Qwen3.8-27B. The model is capable, and the provider economics are easier to live with once you move from testing to real traffic.

Scenario 1: Long-Context RAG Assistant for Internal Docs

A developer-facing assistant over product docs, runbooks, architecture notes, and support history is exactly the kind of workload where Qwen3.8-27B’s 262,144-token context window is useful, and where lower input pricing plus cached-token pricing matters because the app tends to reuse the same system prompt, tool definitions, and retrieval scaffolding.

  • Volume: 1,000 requests/month
  • Pattern: 100,000 input tokens/request, 5,000 output tokens/request
VolumeInput TokensOutput TokensProviderMonthly Cost
1,000 requests/month100,000,0005,000,000DeepInfra$24.38

Cost breakdown: 100M input × $0.15/1M = $15.00; 5M output × $1.875/1M = $9.375. Total monthly cost: $24.38. The same workload on Alibaba’s API ($0.50 input / $3.00 output) would run 100M × $0.50 = $50.00 plus 5M × $3.00 = $15.00, for $65.00 total, versus $24.38 on DeepInfra.

  • Why DeepInfra looks good here: The model’s large context window fits document-heavy retrieval well. DeepInfra’s $0.15/1M input is much easier to justify than Alibaba’s $0.50/1M when every request carries a lot of context, and a stable prompt prefix benefits from the $0.038/1M cached rate.

Scenario 2: Coding Copilot for Repo-Level Bug Fixing

This is a strong use case for Qwen3.8-27B. The benchmarks are unusually solid for a 27B open-weight model: 61.7 on SWE-bench Pro, 73.0 on Terminal Bench 2.1, and 79.0 on QwenSWEBench on DeepInfra’s benchmark page. For code assistants, patch generators, or issue triage agents, DeepInfra provides that capability without first-party token rates. For a nearby reference point, the Qwen3.5-27B demo is a useful earlier-generation comparison for reasoning and coding workloads.

  • Volume: 5,000 requests/month
  • Pattern: 20,000 input tokens/request, 8,000 output tokens/request
VolumeInput TokensOutput TokensProviderMonthly Cost
5,000 requests/month100,000,00040,000,000DeepInfra$90.00

Cost breakdown: 100M input × $0.15/1M = $15.00; 40M output × $1.875/1M = $75.00. Total monthly cost: $90.00. The same workload on Alibaba’s API runs 100M × $0.50 = $50.00 plus 40M × $3.00 = $120.00, for $170.00, versus $90.00 on DeepInfra.

  • Why DeepInfra looks good here: Qwen3.8-27B is described as very verbose, so output pricing matters more than usual. DeepInfra’s $1.875/1M output is materially easier to absorb than Alibaba’s $3.00/1M when the model generates long code patches, explanations, or JSON tool payloads. Function calling and JSON output support help if the copilot sits inside an automated repair or CI workflow.

Scenario 3: Multimodal Browser Agent for Support and Ops Workflows

An agent that reads screenshots, inspects web pages, calls tools, and returns structured actions runs cleanly on Qwen3.8-27B through DeepInfra. The provider exposes image and video input, function calling, and JSON output, and its benchmark page shows strong multimodal agent results including 84.3 on OSWorld-Verified, 64.8 on WebArena-Verified, and 62.9 on Vision2Web.

  • Volume: 2,000 requests/month
  • Pattern: 30,000 input tokens/request, 6,000 output tokens/request
VolumeInput TokensOutput TokensProviderMonthly Cost
2,000 requests/month60,000,00012,000,000DeepInfra$31.50

Cost breakdown: 60M input × $0.15/1M = $9.00; 12M output × $1.875/1M = $22.50. Total monthly cost: $31.50. The same workload on Alibaba’s API runs 60M × $0.50 = $30.00 plus 12M × $3.00 = $36.00, for $66.00, versus $31.50 on DeepInfra.

  • Why DeepInfra looks good here: Multimodal agents accumulate cost from several places at once: tool schemas, screenshots, browser state, and long action traces. DeepInfra is cheaper on both input and output, not just one side of the bill. Private endpoint availability is also relevant for a controlled production deployment, and if the workflow includes image preprocessing, DeepInfra’s background removal model can run on the same platform.

Scenario 4: High-Reuse Agent with Stable Prompts and Cacheable Scaffolding

Some applications are not especially large per request but repeat a lot of the same prompt structure: system instructions, tool definitions, formatting rules, and planning policies. This is where DeepInfra’s explicit cached-token rate of $0.038/1M becomes especially interesting for developers deliberate about prompt design.

  • Volume: 10,000 requests/month
  • Pattern: 10,000 cached input tokens/request, 5,000 fresh input tokens/request, 2,000 output tokens/request
VolumeInput TokensOutput TokensProviderMonthly Cost
10,000 requests/month50M fresh + 100M cached20,000,000DeepInfra$48.80

Cost breakdown: 50M fresh input × $0.15/1M = $7.50; 100M cached input × $0.038/1M = $3.80; 20M output × $1.875/1M = $37.50. Total monthly cost: $48.80. Treating the same 150M total input tokens at Alibaba’s uncached $0.50 rate runs 150M × $0.50 = $75.00 plus 20M × $3.00 = $60.00, for $135.00, versus $48.80 on DeepInfra.

  • Why DeepInfra looks good here: This is the most provider-sensitive pattern in the list. A reusable prompt prefix gets a clearly stated cached-token rate rather than implicit cache economics, which makes Qwen3.8-27B considerably more practical for agent frameworks with repeated schemas and fixed orchestration prompts.

Conclusion

The provider decision for Qwen3.8-27B is about how much you pay per token for the same weights running on someone else’s hardware. Alibaba’s API gives first-party access, which matters if your organization has compliance or contractual reasons to stay close to the model developer. For most production workloads, the pricing gap is wide enough to change the unit economics of the whole application. It is not just a line item in a budget spreadsheet.

Three things are worth weighing carefully: output pricing, cache behavior, and multimodal input cost. Qwen3.8-27B is verbose by design. Long reasoning traces, detailed code responses, and structured tool outputs are part of what makes it useful, and they are also what makes output token pricing the dominant cost driver in most real workloads. A stable system prompt or repeating scaffolding makes DeepInfra’s cached rate of $0.038 per 1M a real lever, not a footnote. For multimodal workflows, images and video frames still count as input tokens, and the multimodal model options on DeepInfra show what else is available alongside Qwen3.8-27B.

If you want to see how the model behaves before committing to a workload estimate, the Qwen3.8-27B demo on DeepInfra is a practical starting point: run a few representative prompts, check output length, and get a feel for latency before building pricing assumptions into your architecture. When you are ready to integrate, the API documentation covers authentication, request format, function calling, and multimodal input handling. The model is capable. The job is making sure your cost model reflects how it actually behaves in your use case.

Related articles
GLM-5.1 Pricing Guide: API Cost Comparison & AnalysisGLM-5.1 Pricing Guide: API Cost Comparison & Analysis<p>Provider choice for GLM-5.1 is a real economic decision. Across 10 benchmarked API providers, blended pricing runs from $0.74 to $1.70 per 1M tokens, output speed from 33.8 to 175.2 t/s, and the fastest provider is 5.2x quicker than the slowest. For teams deploying at scale, that spread determines whether this model fits a production [&hellip;]</p>
Top 6 GLM-5.2 Max API Providers ComparedTop 6 GLM-5.2 Max API Providers Compared<p>Deploying the GLM-5.2 (max) Mixture-of-Experts model — 753B total parameters with roughly 40B active per token and a 1M context window — requires infrastructure that separates production-grade API providers from the rest. This guide breaks down the top providers by throughput, latency, pricing, and quantization architecture. GLM-5.2 (max) API Review Summary (2026-06-27) TL;DR: Best Providers [&hellip;]</p>
Agents Need a Runtime Boundary, Not Just GuardrailsAgents Need a Runtime Boundary, Not Just GuardrailsAs agents move beyond model calls to tools, APIs, and execution environments, the runtime becomes part of the architecture. How DeepInfra Sandboxes and Hosted Agents keep the boundary outside the agent, where NVIDIA OpenShell fits, and why we plan around NVIDIA Vera.