DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Qwen3.8-27B is Alibaba’s 27B open-weight reasoning model, released August 14, 2026 under Apache 2.0. It ships with a large context window, multimodal input, and API availability across eight providers. Pricing varies sharply by host: Alibaba’s first-party rate runs over 3x DeepInfra’s for the same weights.
Context window figures differ slightly by source. Artificial Analysis reports 256K tokens; DeepInfra lists a native 262,144, extensible to 1,000,000. Both describe the same model: text, image, and video input, text-only output, and public weights on Hugging Face, with JSON output, function calling, and multimodal workflows supported.
Artificial Analysis gives it an Intelligence Index score of 34, against a median of 8 for comparable open-weight models, well above average in its class. DeepInfra’s benchmark page shows the same strength against Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B, and Opus4.6 Max. It scores 73.0 on Terminal Bench 2.1, 61.7 on SWE-bench Pro, 84.3 on OSWorld-Verified, and 64.8 on WebArena-Verified. The tradeoff is efficiency. Artificial Analysis classifies the model as notably slow and very verbose: 44.8 output tokens per second, 3.88 seconds to first token, and 200M output tokens generated during Intelligence Index evaluation.
This verbosity makes provider choice matter more than usual. Alibaba’s API prices Qwen3.8-27B at $0.50 per 1M input tokens and $3.00 per 1M output tokens, rates Artificial Analysis calls expensive relative to similar open-weight models. DeepInfra lists the same model at $0.15 input, $1.875 output, and $0.038 cached per 1M, a meaningfully lower bill for long-context RAG, multimodal agents, and sustained production traffic.
Qwen3.8-27B looks strongest when you want a high-performing open-weight reasoning model and care about deployment flexibility as much as raw benchmarks. The pricing spread is wide. Alibaba’s API is listed at $0.50 input / $3.00 output per 1M tokens, while DeepInfra lists $0.15 input / $1.875 output / $0.038 cached. The best provider choice depends less on the model itself than on whether you prioritize first-party access or lower-cost production inference. Compared with Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B, and Opus4.6 Max, Qwen3.8-27B is a serious contender for coding, agent, and multimodal workloads, but verbosity and latency are worth watching.
| Best For | Provider | Why |
|---|---|---|
| Lowest price / cost-sensitive workloads | DeepInfra | DeepInfra’s standard pricing of $0.15 input and $1.875 output per 1M tokens is lower than Alibaba’s $0.50 input and $3.00 output for the same model. |
| RAG, document-heavy, or high-throughput use cases | DeepInfra | DeepInfra combines a 262,144-token public context window with cached-token pricing at $0.038 per 1M, which fits long-context and repeat-query workloads. |
| Easiest onboarding / fastest time-to-first-call | DeepInfra | The model is deployment-ready on DeepInfra’s inference cloud with public endpoint access, JSON output, function calling, multimodal input, and private endpoint availability. |
| Multimodal agents and tool-using applications | DeepInfra | DeepInfra exposes image and video input plus function calling, with strong results on OSWorld-Verified, WebArena-Verified, AndroidWorld, and Vision2Web. |
| First-party managed access | Alibaba API | Artificial Analysis sources its pricing and latency metrics directly from Alibaba’s first-party API, the clearest reference point for official hosted access. |
| Benchmark-first evaluation against nearby alternatives | DeepInfra | DeepInfra provides side-by-side benchmark comparisons versus Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B, and Opus4.6 Max across coding, agent, general, and multimodal tasks. |
Tokens are the billing unit for LLM APIs, and they are not the same thing as words. A token can be a whole short word, part of a long word, punctuation, JSON syntax, a code symbol, or a whitespace pattern. For practical budgeting: 1M tokens is a lot of text. Production bills usually come from repetition, long prompts, agent loops, and verbose outputs, more than from a single big request.
| Token Type | What It Is | Why It Matters |
|---|---|---|
| Input tokens | Everything you send to the model: system prompt, user prompt, chat history, tool schemas, retrieved docs, and images or video encoded into the request. | This is the part that quietly grows over time. RAG, long chat threads, and large tool definitions can make input cost dominate faster than expected. |
| Output tokens | Everything the model generates back: answers, JSON, code, tool-call arguments, and long-form completions. | Qwen3.8-27B is very verbose in evaluation. If your app allows long answers, output cost can become the main line item. |
| Cached tokens | Reused prompt tokens the provider recognizes from previous requests and bills at a reduced rate. | This matters most for stable prefixes: system prompts, repeated instructions, and recurring RAG scaffolding. Cache pricing can materially change total cost. |
Token accounting matters more than usual for Qwen3.8-27B because the model combines a very large context window, multimodal input, reasoning support, and a tendency toward long outputs. It is a useful mix for agents and document-heavy workflows. It is also how a bill that looked fine in a spreadsheet looks less fine in production. A few cost behaviors worth keeping in mind:
The headline difference is simple: the same model is much cheaper on DeepInfra than on Alibaba’s first-party API. The less obvious difference is where the savings show up: DeepInfra is lower on input tokens, lower on output tokens, and exposes a clearly listed cached-token rate. If you run long-context RAG, repeat a large system prompt, or keep agent instructions stable across calls, cache pricing matters more than most teams assume at the start.
| Provider | Input Tokens | Output Tokens | Cached Tokens | Notes |
|---|---|---|---|---|
| DeepInfra | $0.15 / 1M | $1.875 / 1M | $0.038 / 1M | Lowest listed cost in this comparison. Best fit for prompt-heavy apps, long-context RAG, and repeated workflows where cache hits are common. |
| Alibaba API | $0.50 / 1M | $3.00 / 1M | 90% cache discount off list | Clear first-party reference pricing. Much more expensive on both input and output, and output pricing is especially costly for a verbose model. |
The provider choice changes the economics quickly:
The practical read: if you care about token efficiency, DeepInfra is the easier default for Qwen3.8-27B. If you want first-party managed access, Alibaba is the reference option, but you pay a real premium for the same model’s tokens. If your app has high cache reuse, test both with real prompts, since cache behavior can change total cost more than raw input price suggests. If your app lets the model think out loud or produce long JSON and code responses, cap output length early, or Qwen3.8-27B will spend your budget for you.
DeepInfra runs inference on bare-metal infrastructure, which matters because cutting out extra virtualization layers helps with both performance consistency and cost efficiency. It is typically 50 to 80% cheaper than major cloud competitors, with appeal among developers, high-volume API users, and teams watching unit economics closely. For teams that want strong open-model capability without premium-hosting prices, browsing the full model catalog is a practical place to start.
DeepInfra lists Qwen3.8-27B at $0.15 per 1M input tokens and $1.875 per 1M output tokens, against $0.50 input and $3.00 output on Alibaba’s API for the same model, a meaningful reduction on both sides of the bill, especially for prompt-heavy workloads or long responses. If you expect sustained traffic, the provider choice directly changes whether this model feels affordable in production.
Below are practical scenarios where DeepInfra is a particularly strong fit for Qwen3.8-27B. The model is capable, and the provider economics are easier to live with once you move from testing to real traffic.
A developer-facing assistant over product docs, runbooks, architecture notes, and support history is exactly the kind of workload where Qwen3.8-27B’s 262,144-token context window is useful, and where lower input pricing plus cached-token pricing matters because the app tends to reuse the same system prompt, tool definitions, and retrieval scaffolding.
| Volume | Input Tokens | Output Tokens | Provider | Monthly Cost |
|---|---|---|---|---|
| 1,000 requests/month | 100,000,000 | 5,000,000 | DeepInfra | $24.38 |
Cost breakdown: 100M input × $0.15/1M = $15.00; 5M output × $1.875/1M = $9.375. Total monthly cost: $24.38. The same workload on Alibaba’s API ($0.50 input / $3.00 output) would run 100M × $0.50 = $50.00 plus 5M × $3.00 = $15.00, for $65.00 total, versus $24.38 on DeepInfra.
This is a strong use case for Qwen3.8-27B. The benchmarks are unusually solid for a 27B open-weight model: 61.7 on SWE-bench Pro, 73.0 on Terminal Bench 2.1, and 79.0 on QwenSWEBench on DeepInfra’s benchmark page. For code assistants, patch generators, or issue triage agents, DeepInfra provides that capability without first-party token rates. For a nearby reference point, the Qwen3.5-27B demo is a useful earlier-generation comparison for reasoning and coding workloads.
| Volume | Input Tokens | Output Tokens | Provider | Monthly Cost |
|---|---|---|---|---|
| 5,000 requests/month | 100,000,000 | 40,000,000 | DeepInfra | $90.00 |
Cost breakdown: 100M input × $0.15/1M = $15.00; 40M output × $1.875/1M = $75.00. Total monthly cost: $90.00. The same workload on Alibaba’s API runs 100M × $0.50 = $50.00 plus 40M × $3.00 = $120.00, for $170.00, versus $90.00 on DeepInfra.
An agent that reads screenshots, inspects web pages, calls tools, and returns structured actions runs cleanly on Qwen3.8-27B through DeepInfra. The provider exposes image and video input, function calling, and JSON output, and its benchmark page shows strong multimodal agent results including 84.3 on OSWorld-Verified, 64.8 on WebArena-Verified, and 62.9 on Vision2Web.
| Volume | Input Tokens | Output Tokens | Provider | Monthly Cost |
|---|---|---|---|---|
| 2,000 requests/month | 60,000,000 | 12,000,000 | DeepInfra | $31.50 |
Cost breakdown: 60M input × $0.15/1M = $9.00; 12M output × $1.875/1M = $22.50. Total monthly cost: $31.50. The same workload on Alibaba’s API runs 60M × $0.50 = $30.00 plus 12M × $3.00 = $36.00, for $66.00, versus $31.50 on DeepInfra.
Some applications are not especially large per request but repeat a lot of the same prompt structure: system instructions, tool definitions, formatting rules, and planning policies. This is where DeepInfra’s explicit cached-token rate of $0.038/1M becomes especially interesting for developers deliberate about prompt design.
| Volume | Input Tokens | Output Tokens | Provider | Monthly Cost |
|---|---|---|---|---|
| 10,000 requests/month | 50M fresh + 100M cached | 20,000,000 | DeepInfra | $48.80 |
Cost breakdown: 50M fresh input × $0.15/1M = $7.50; 100M cached input × $0.038/1M = $3.80; 20M output × $1.875/1M = $37.50. Total monthly cost: $48.80. Treating the same 150M total input tokens at Alibaba’s uncached $0.50 rate runs 150M × $0.50 = $75.00 plus 20M × $3.00 = $60.00, for $135.00, versus $48.80 on DeepInfra.
The provider decision for Qwen3.8-27B is about how much you pay per token for the same weights running on someone else’s hardware. Alibaba’s API gives first-party access, which matters if your organization has compliance or contractual reasons to stay close to the model developer. For most production workloads, the pricing gap is wide enough to change the unit economics of the whole application. It is not just a line item in a budget spreadsheet.
Three things are worth weighing carefully: output pricing, cache behavior, and multimodal input cost. Qwen3.8-27B is verbose by design. Long reasoning traces, detailed code responses, and structured tool outputs are part of what makes it useful, and they are also what makes output token pricing the dominant cost driver in most real workloads. A stable system prompt or repeating scaffolding makes DeepInfra’s cached rate of $0.038 per 1M a real lever, not a footnote. For multimodal workflows, images and video frames still count as input tokens, and the multimodal model options on DeepInfra show what else is available alongside Qwen3.8-27B.
If you want to see how the model behaves before committing to a workload estimate, the Qwen3.8-27B demo on DeepInfra is a practical starting point: run a few representative prompts, check output length, and get a feel for latency before building pricing assumptions into your architecture. When you are ready to integrate, the API documentation covers authentication, request format, function calling, and multimodal input handling. The model is capable. The job is making sure your cost model reflects how it actually behaves in your use case.
GLM-5.1 Pricing Guide: API Cost Comparison & Analysis<p>Provider choice for GLM-5.1 is a real economic decision. Across 10 benchmarked API providers, blended pricing runs from $0.74 to $1.70 per 1M tokens, output speed from 33.8 to 175.2 t/s, and the fastest provider is 5.2x quicker than the slowest. For teams deploying at scale, that spread determines whether this model fits a production […]</p>
Top 6 GLM-5.2 Max API Providers Compared<p>Deploying the GLM-5.2 (max) Mixture-of-Experts model — 753B total parameters with roughly 40B active per token and a 1M context window — requires infrastructure that separates production-grade API providers from the rest. This guide breaks down the top providers by throughput, latency, pricing, and quantization architecture. GLM-5.2 (max) API Review Summary (2026-06-27) TL;DR: Best Providers […]</p>
Agents Need a Runtime Boundary, Not Just GuardrailsAs agents move beyond model calls to tools, APIs, and execution environments, the runtime becomes part of the architecture. How DeepInfra Sandboxes and Hosted Agents keep the boundary outside the agent, where NVIDIA OpenShell fits, and why we plan around NVIDIA Vera.© 2026 DeepInfra. All rights reserved.