DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Qwen3.8-27B landed at #9 on Arena.ai’s Code Arena WebDev board with 1,595 points, and as of September 2026 it’s the only model under 30 billion parameters in the top 10. Gemma 4-31B, released the same year and larger by three billion parameters, sits at #80. Three ranks above the 27B sits Qwen3.8-2.4T-A95B, a sibling with 86 times its parameter count (the A95B suffix is the 95 billion it actually runs per token; the rest of this article is about why that distinction sets the price).
Most coverage of open models tracks one gap, how close the biggest open weights get to GPT and Claude. Small open weight models have been closing a second one that gets far less attention: the distance between what fits on one GPU and what needs a rack.
A smaller, more efficient model should cost less to run. Then you open the rate card. On DeepInfra, Qwen3.8-27B costs $3.00 per million output tokens. Gemma 4-31B costs $0.38, and it’s the bigger model. Per-token price is the wrong place to look for the efficiency gain.
Epoch AI puts the lag between the best open-weight models and frontier closed models at four months on its Capabilities Index since January 2026. The UK AI Safety Institute measured 4 to 7 months on cyber capability, down from the 6 to 10 months it saw across open models released in 2025. That story is well covered, and DeepInfra has written its own read of the open versus closed gap on intelligence, price, and speed.
The other axis gets less attention. Hugging Face’s state of open models report put China’s monthly parameter ceiling between 754 billion and 2.78 trillion, while US models stayed under 130 billion in five of seven months. Alibaba works both ends of that range at once. It ships frontier-scale weights, and it ships a 27B that outranks most of the field on agentic coding. The same report counts 151,448 Qwen-derived repositories on the Hub, growing by 180 to 210 a day, which is what a size band with real demand behind it looks like.
If you serve inference rather than read leaderboards, the second gap is the one with money attached. A four-month capability lag decides which model you pick. A 27B model at #9 decides which hardware bill you sign.
Here is the live catalog, sorted by output price. Active parameters are the ones read for every token emitted. In a dense model that means all of them. In a mixture-of-experts (MoE) model, a small gating layer called the router picks a handful of expert blocks per token, and the active count is whatever it picks. Served precision is what DeepInfra runs, which is not always what the lab published.
| Model | Total / active params | Context | Input / 1M | Output / 1M | Served precision |
|---|---|---|---|---|---|
| gpt-oss-120b | 117B / 5.1B | 131K | $0.04 | $0.17 | bfloat16 |
| Nemotron-3-Nano-30B-A3B | 30B / 3B | 262K | $0.05 | $0.20 | fp4 |
| Mistral-Small-3.2-24B | 24B dense | 128K | $0.07 | $0.20 | fp8 |
| gemma-4-26B-A4B-it | 26B / 4B | 262K | $0.07 | $0.34 | fp8 |
| gemma-4-31B-it | 31B dense | 262K | $0.13 | $0.38 | fp8 |
| Qwen3.6-35B-A3B | 35B / 3B | 262K | $0.10 | $0.95 | fp8 |
| Qwen3.5-122B-A10B | 122B / 10B | 262K | $0.29 | $2.40 | fp4 |
| Qwen3.8-27B | 27.8B dense | 262K | $0.40 | $3.00 | none listed |
| Qwen3.5-397B-A17B | 397B / 17B | 262K | $0.45 | $3.00 | fp8 |
| Qwen3.8-2.4T-A95B | 2.4T / 95B | 262K | $2.00 | $6.00 | fp4 |
Source: live DeepInfra catalog, September 4, 2026.
Sort by price and the size ordering falls apart. gpt-oss-120b carries four times the parameters of Qwen3.8-27B at a seventeenth the output price. Qwen3.8-27B is priced identically to Qwen3.5-397B-A17B, fourteen times its size. The sharpest pair sits in the middle. Gemma 4-31B is dense, bigger, and 7.9 times cheaper per output token than the 27B that outranks it by 71 places on Code Arena, and its full pricing and benchmark breakdown walks the rest of that family.
Shopping by parameter count means reading a column that doesn’t predict the bill.
Emitting a token means reading the active weights out of GPU memory, and that read dominates decode cost. So the first dial is active parameters, not total ones. gpt-oss-120b touches 5.1 billion of its 117 billion per token, a figure DeepInfra publishes on the model page. Qwen3.8-27B is dense, so it touches all 27.8 billion, roughly five and a half times the memory traffic for the same forward pass.
The other dial is served precision, which DeepInfra also publishes per model. Gemma 4 ships in three variants off the same weights: fp4 turbo at $0.09 in and $0.34 out, fp8 at $0.13 and $0.38, and an fp8 Ultra tier at $0.27 and $0.76. Fewer bytes per parameter means less to move, and the rate card reflects it.
Both dials are readable from the catalog itself:
import json, urllib.request
CATALOG = "https://api.deepinfra.com/models/list"
WATCH = {
"Qwen/Qwen3.8-27B",
"google/gemma-4-31B-it",
"google/gemma-4-31B-it-turbo",
"openai/gpt-oss-120b",
}
with urllib.request.urlopen(CATALOG) as resp:
catalog = json.load(resp)
for m in catalog:
if m["model_name"] not in WATCH:
continue
p = m["pricing"]
dollars_in = p["cents_per_input_token"] * 10_000
dollars_out = p["cents_per_output_token"] * 10_000
cached = dollars_in * (p["rate_per_input_token_cached"] or 1.0)
print(
f"{m['model_name']:30} {m['quantization'] or 'unquantized':12}"
f" in ${dollars_in:.3f} cached ${cached:.3f} out ${dollars_out:.2f}"
)The Qwen3.8-27B row comes back with no quantization listed. DeepInfra doesn’t publish a served precision for this one, so the honest assumption is that you’re getting something close to the weights Alibaba released, the same file that scored on the leaderboard. That’s an assumption, not a spec.
Together the dials explain a floor, not the whole spread. At assumed bf16, Qwen3.8-27B moves about 56 GB of weights per token pass against Gemma 4-31B’s 31 GB at fp8. That’s a 1.8x hardware gap sitting under a 7.9x price gap. The rest is demand, workload shape, and margin. Reasoning models bill in the output lane, and a model pointed at long agent runs emits far more output tokens per request than a general instruct model does. The rate card is a price list, not an efficiency meter.
Alibaba has now filled the same size slot three generations running, at the same 262K context, on the same endpoint.
| Model | Live on DeepInfra | Input / 1M | Output / 1M | Cached input | Served precision |
|---|---|---|---|---|---|
| Qwen3.5-27B | 2026-03-24 | $0.26 | $2.60 | n/a | fp8 |
| Qwen3.6-27B | 2026-04-30 | $0.32 | $3.20 | n/a | fp8 |
| Qwen3.8-27B | 2026-08-17 | $0.40 | $3.00 | $0.04 | none listed |
Input climbed 54 percent across five months. If efficiency arrived as a lower per-token price, this is the row where it would show up, and it doesn’t. DeepInfra’s Qwen3.5 27B benchmark run clocked that generation at 153.3 output tokens per second and $0.84 per million blended across its input-output mix, so the older model is still the cheaper way to move a token.
What the extra 14 cents buys is a different model wearing the same parameter count: vision input, 262K of hybrid attention, and thinking you can switch off. Hybrid attention means most layers use linear attention, with full-attention layers at intervals, so the KV cache that usually dominates long-context memory stays small. The catalog tags Qwen3.8-27B can-disable-reasoning, and DeepInfra’s reasoning docs are blunt about why that matters, noting that reasoning tokens count toward output token billing. At $3.00 per million, the first cost control on this model is deciding per request whether you need the thinking at all.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["DEEPINFRA_API_TOKEN"],
base_url="https://api.deepinfra.com/v1/openai",
)
# Ticket triage does not need a chain of thought. Switching reasoning off
# keeps the output lane, the expensive one, from carrying tokens you pay
# for and never read.
resp = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=[
{"role": "system", "content": "Label the ticket: bug, billing, or feature."},
{"role": "user", "content": ticket_text},
],
extra_body={"reasoning_effort": "none"},
max_tokens=8,
)
print(resp.choices[0].message.content, resp.usage.completion_tokens)Engram and Harvey built a synthetic law firm, 9,286 files and 266 client matters across 100 million tokens, then adapted Qwen3.8-27B to it with parametric memory (domain knowledge baked into the weights, rather than fetched at query time) and retrieval. On strict all-or-nothing correctness the adapted 27B beat Claude Opus 4.8, 30 percent to 25, at roughly $0.13 per query against $1.32. Tokens in completed trajectories fell 58 percent.
Read the fine print. That’s a fine-tuned model on a domain corpus rather than a stock endpoint, so part of the gap belongs to the adaptation. The mechanism still transfers. A model that finishes in fewer turns bills fewer tokens, and turns are where agent cost lives.
Price a realistic loop and the shape shows up. Take 12 turns per task, a 6,000-token system and tool prefix that stays stable, 1,500 fresh tokens of conversation per turn, and 700 output tokens per turn. The prefix lands fresh on turn one, then rides the cache for the other eleven turns: that’s 24,000 fresh input tokens, 66,000 cache-eligible tokens, and 8,400 output tokens per task. This workload deliberately caches only the static prefix; a real agent also accumulates conversation history, which pushes the cache share higher, so treat these tables as the conservative case.
| Model | Cached input | Fresh input | Output | Per 1,000 tasks |
|---|---|---|---|---|
| gemma-4-31B-it | $8.58 | $3.12 | $3.19 | $14.89 |
| Qwen3.8-27B | $2.64 | $9.60 | $25.20 | $37.44 |
| Qwen3.8-2.4T-A95B | $13.20 | $48.00 | $50.40 | $111.60 |
Estimate, using live DeepInfra rates and the workload above. Your turn count is the variable that matters.
Two things fall out. Output is the lane that decides the bill on a reasoning model, and caching only ever touches the other one. And the 27B runs the same trajectory for a third of what its 2.4T sibling charges, six ranks apart on the coding board. Whether the cheaper rung closes your tasks in one turn is a measurement, not an assumption, which is what the instrumentation below is for.
Instrument it rather than assuming it:
PRICES = { # USD per million tokens, live DeepInfra rates
"Qwen/Qwen3.8-27B": {"fresh": 0.40, "cached": 0.04, "out": 3.00},
"google/gemma-4-31B-it": {"fresh": 0.13, "cached": 0.13, "out": 0.38},
}
def turn_cost(model: str, usage) -> float:
p = PRICES[model]
details = getattr(usage, "prompt_tokens_details", None)
cached = getattr(details, "cached_tokens", 0) or 0
fresh = usage.prompt_tokens - cached
return (
fresh * p["fresh"] + cached * p["cached"] + usage.completion_tokens * p["out"]
) / 1_000_000
spent, turns = 0.0, 0
for turn in run_agent(task): # one chat.completions call per turn
spent += turn_cost(turn.model, turn.usage)
turns += 1
print(f"{task.id}: {turns} turns, ${spent:.4f}, ${spent / max(turns, 1):.4f} per turn")The Qwen3.8-27B that placed ninth is a 55.58 GB BF16 checkpoint. The weights aren’t the whole bill: only 16 of its 64 layers run full attention, which keeps the KV cache to 16 GiB at the full 262K context, and the practical floor lands at roughly 67.8 GiB, which means an 80GB card or several smaller ones ganged together. The build that fits a 24GB consumer GPU is a Q4_K_M GGUF at 17.11 GB, and plenty of people run smaller than that.
That distinction does quiet work inside the local-inference cost arguments. One widely shared walkthrough clocks a Q3 quant at 41 tokens per second on a consumer card, prices 250W at 21 cents per kilowatt hour, and lands at about 37 cents per million output tokens against a $3.00 endpoint rate. The arithmetic holds. The scope is narrow: electricity only, one stream at a time, on a file about a quarter the size of the one that scored.
Quantization loss is measurable rather than theoretical. The same walkthrough charts Kullback-Leibler divergence, the gap between the quantized model’s next-token distribution and the original’s, against file size across community builds. Q3_K_M sits near 0.07, roughly triple the divergence of the Q4 builds one tier up. That model makes different choices under load, and calling it the same weights at a discount understates what changed.
That’s why the “none listed” cell in the price table is worth a second look. Served precision is a variable you can check per model before you point traffic at it.
Once cost per task is the metric, model choice becomes a per-step decision rather than an application-level one. Two levers do most of the work: context caching and model routing.
Context caching comes first. DeepInfra matches prompt prefixes automatically and reports the hit in the usage object, so the job is ordering the message list to put stable content first. On Qwen3.8-27B that turns $0.40 input into $0.04, and an agent resending its tool schemas every turn is the shape that benefits most.
SYSTEM = TOOL_SCHEMAS + HOUSE_STYLE # stable for the life of the run
resp = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=[
{"role": "system", "content": SYSTEM}, # stable prefix first
*history, # then whatever changed
{"role": "user", "content": step},
],
extra_body={"prompt_cache_key": f"agent-{task.id}"},
)
d = resp.usage.prompt_tokens_details
print(f"{d.cached_tokens}/{resp.usage.prompt_tokens} prompt tokens came from cache")Model routing comes second. Start cheap, validate, climb only on failure. Every rung shares a base URL and a key, so a rung is a string:
LADDER = [
("google/gemma-4-31B-it", {"reasoning_effort": "none"}),
("Qwen/Qwen3.8-27B", {"reasoning_effort": "medium"}),
("Qwen/Qwen3.8-2.4T-A95B", {"reasoning_effort": "high"}),
]
def solve(messages, passes):
for model, opts in LADDER:
resp = client.chat.completions.create(
model=model, messages=messages, extra_body=opts
)
answer = resp.choices[0].message.content
if passes(answer):
return model, answer
raise RuntimeError("no rung produced a valid answer")The ladder pays off only if the cheap rung clears the bar often enough. How often? That’s a per-workload measurement. Log which rung closed each task alongside the per-turn cost from the previous section, and escalation rate becomes a number you can act on. If the top rung closes most tasks, the ladder is buying latency and nothing else. Start there instead. Swapping a rung costs one string either way, because all three models sit behind the same OpenAI-compatible endpoint. DeepInfra’s own breakdown of routing and caching savings walks the same pattern for coding agents, and the Qwen rate card prices the rest of the family you might put on the rungs.
On coding and agentic work, yes. Qwen3.8-27B reports LiveCodeBench v6 at 90.3 against Gemma 4-31B’s 80.0, and the #9 versus #80 split is a WebDev coding board, not a general ranking. Gemma 4-31B scores higher on AIME 2026 and MMLU, and costs a fraction as much per output token. Pick per workload.
Yes, at reduced precision and reduced context. The Q4_K_M GGUF is 17.11 GB, which leaves room for a cache at 32K to 64K tokens on a 24GB card. Full BF16 weights alone take 55.58 GB, and the KV cache at the native 262K context adds 16 GiB, pushing the requirement to about 67.8 GiB.
It changes what the model is. DeepInfra’s docs note that disabling reasoning makes a reasoning model behave like a standard chat model. For classification, extraction, and routing that’s usually the behavior you want and the cheaper bill is free. For multi-step planning and repository-scale coding, thinking is where Qwen3.8-27B’s reported gains come from.
No. Prefix caching is automatic on DeepInfra and needs no extra parameters. Structure prompts so stable content leads, then read prompt_tokens_details.cached_tokens off the usage object to confirm hits. The optional prompt_cache_key parameter improves hit rates across similar prompts.
Apache 2.0, which permits commercial use and self-hosting. That’s the norm rather than the exception in this size class now. Of 178 Chinese releases above 20B parameters this year, 59 percent carry Apache 2.0 and 22 percent MIT.
Small open weight models closed a real gap this year, and the rate card is the wrong instrument for measuring it. Per-token price sets a floor with active parameters and served precision; demand and workload decide everything above it. What fell is the number of tokens it takes to finish a job, and that shows up only when you meter tasks instead of tokens.
Running the whole size range behind one OpenAI-compatible endpoint is what makes that measurable, because comparing two rungs costs a string change. Start on the Qwen3.8-27B model page for live pricing, or browse the open-weight catalog for the rest of the ladder. Tell us what your escalation rate looks like at feedback@deepinfra.com, on Discord, or at @DeepInfra.
vLLM vs SGLang: Performance, Features & Deployment Compared<p>Somebody on your team read a benchmark post, and now there’s a ticket to migrate the inference stack. That’s how most vLLM vs SGLang decisions start. A published test reports a 29 percent throughput gap, the number lands in Slack, and two weeks later you’re debugging kernel version conflicts at midnight while p99 latency sits […]</p>
Open-Source vs Closed-Source AI Models: Is the Gap Worth It?<p>The Artificial Analysis Intelligence Index sits at a ceiling of 57. Three frontier models — Claude Opus 4.7, Gemini 3.1 Pro Preview, and GPT-5.5 — all land in that band. Meanwhile, four open-weight models released between February and April 2026 now score 50 or above on the same index. A year ago, the best open-weight […]</p>
© 2026 DeepInfra. All rights reserved.