DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

The Qwen Team’s latest open-weight release, Qwen3.8-27B, is a 27-billion-parameter vision-language model. It goes deep on the tasks that trip up most models: multi-step agentic workflows, software engineering, and multimodal reasoning across images, documents, and video. On SWE-bench Pro, it scores 61.7, outpacing Opus4.6 Max’s 53.4. On QwenSWEBench it jumps from 49.3 on the previous generation to 79.0, a gap that’s hard to read as incremental progress.
What makes Qwen3.8-27B particularly interesting for production use is its hybrid architecture: 64 layers mixing Gated DeltaNet linear attention with standard gated attention. It pairs that with a native context window of 262,144 tokens, extensible to one million via YaRN RoPE scaling. Thinking mode is on by default but fully controllable per request, with tunable reasoning depth via reasoning_effort. You can trade deep deliberation for fast response without swapping models. The vision-language benchmarks reflect the same pattern: OSWorld-Verified hits 84.3, ahead of both Qwen3.6-27B at 63.9 and Qwen3.7-Plus at 73.3.
The model is now available on DeepInfra.
Qwen3.8-27B is a 27B-parameter dense vision-language model, but its architecture is not a standard transformer. It uses a hybrid linear and full attention design across 64 layers: 16 repeating blocks of 3 x (Gated DeltaNet to FFN) followed by 1 x (Gated Attention to FFN). Three-quarters of the layers use linear attention, Gated DeltaNet with 48 heads for V and 16 for QK. One quarter uses standard gated attention, 24 Q heads, 4 KV heads, head dimension 256. The layout balances long-context efficiency with the expressivity of full attention, rather than defaulting to one or the other. Native context length is 262,144 tokens, extensible to 1M via YaRN RoPE scaling.
The gains over Qwen3.5-27B and the previous Qwen3.6-27B are substantial, particularly in coding and agentic tasks:
| Benchmark | Qwen3.8-27B | Qwen3.6-27B | Change |
|---|---|---|---|
| SWE-bench Pro | 61.7 | 53.5 | +8.2 |
| DeepSWE 1.1 | 42.2 | 13.3 | +28.9 |
| QwenSWEBench | 79.0 | 49.3 | +29.7 |
| LiveCodeBench v6 | 90.3 | 83.9 | +6.4 |
| OSWorld-Verified | 84.3 | 63.9 | +20.4 |
| IFBench | 79.5 | 69.1 | +10.4 |
The DeepSWE and QwenSWEBench jumps (+28.9 and +29.7 respectively) stand out. These are real-world software engineering task benchmarks, not synthetic reasoning evals.
On the multimodal side, Qwen3.8-27B natively handles images and videos as input, including STEM diagrams, documents, and hour-scale videos. The vision benchmarks reflect genuine capability here: 84.3 on OSWorld-Verified (versus 72.7 for Opus4.6 Max), 81.9 on AndroidWorld, and 90.0 on MathVision without code interpreter. These scores put it ahead of larger closed models on several GUI and visual agent tasks. You can explore its multimodal capabilities alongside other multimodal models on DeepInfra.
One of the more practically useful features is flexible thinking control. Thinking mode is on by default, but can be toggled per request. Reasoning depth is adjustable via reasoning_effort (xhigh, medium, low), and reasoning context from prior turns can be preserved with preserve_thinking. This gives you explicit control over the latency and quality tradeoff at inference time without needing separate model checkpoints.
The model ships in BF16/Safetensors format and is compatible with vLLM, SGLang, TokenSpeed, and Hugging Face Transformers. With a large and growing set of community quantizations already available on Hugging Face, adapting it for lower-memory deployments is straightforward.
Qwen3.8-27B is available on DeepInfra as a public API endpoint. Pricing is usage-based: input tokens run at $0.15 / 1M tokens, output at $1.875 / 1M tokens, and cached tokens at $0.038 / 1M tokens. Priority and Flex tiers are also available on the platform. The 262,144-token context window, function calling, JSON output, and multimodal (image and video) inputs are all supported out of the box. Private endpoint deployment is available through the DeepInfra dashboard if you need dedicated capacity.
DeepInfra exposes Qwen3.8-27B through an OpenAI-compatible API: no infrastructure to provision, no model weights to manage, and no setup beyond grabbing an API key. The full API documentation covers all available parameters, including the reasoning_effort and preserve_thinking fields. DeepInfra operates with a zero data retention policy and is both SOC 2 and ISO 27001 certified.
Here’s how to make your first call:
from openai import OpenAI
client = OpenAI(
api_key="$DEEPINFRA_TOKEN",
base_url="https://api.deepinfra.com/v1/openai",
)
response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=[{"role": "user", "content": "Explain the difference between linear attention and standard attention."}],
)
print(response.choices[0].message.content)
import OpenAI from "openai";
const openai = new OpenAI({
apiKey: "$DEEPINFRA_TOKEN",
baseURL: "https://api.deepinfra.com/v1/openai",
});
const response = await openai.chat.completions.create({
model: "Qwen/Qwen3.8-27B",
messages: [{ role: "user", content: "Explain the difference between linear attention and standard attention." }],
});
console.log(response.choices[0].message.content);
curl "https://api.deepinfra.com/v1/openai/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $DEEPINFRA_TOKEN" \
-d '{
"model": "Qwen/Qwen3.8-27B",
"messages": [
{
"role": "user",
"content": "Explain the difference between linear attention and standard attention."
}
]
}'The only things to swap relative to any existing OpenAI client setup are the base URL (https://api.deepinfra.com/v1/openai), your DeepInfra token, and the model name. Everything else, streaming, tool calling, structured output, works the same way.
If you want to test the model before wiring it into a pipeline, the live demo on the Qwen3.8-27B model page lets you run it directly in the browser. For a broader look at what’s available, the full model catalogue is a useful reference when evaluating alternatives or building multi-model pipelines. If you’re comparing against the previous generation, the Qwen3.5-27B API documentation is worth a look for context on parameter structure and supported features.
Qwen3.8-27B represents a meaningful step forward for open-weight models in the areas where progress is typically slowest: real software engineering tasks, long-context multimodal reasoning, and agentic workflows that demand reliability and flexibility. The hybrid attention architecture and tunable reasoning depth mean you’re not locked into a single performance profile. You can push for maximum quality on complex tasks, or trade reasoning depth for lower latency, within the same model.
It makes Qwen3.8-27B a reasonable starting point for autonomous coding agents, document-grounded pipelines, and GUI automation systems that previously required closed, larger-scale models. If that fits what you’re building, the Qwen3.8-27B demo is a quick way to get a feel for the model. The broader Qwen model family is worth exploring for production workloads.
GLM-5.1 Pricing Guide: API Cost Comparison & Analysis<p>Provider choice for GLM-5.1 is a real economic decision. Across 10 benchmarked API providers, blended pricing runs from $0.74 to $1.70 per 1M tokens, output speed from 33.8 to 175.2 t/s, and the fastest provider is 5.2x quicker than the slowest. For teams deploying at scale, that spread determines whether this model fits a production […]</p>
Unleashing the Potential of AI for Exceptional Gaming ExperiencesGaming companies are constantly in search of ways to enhance player experiences and achieve
extraordinary outcomes. Recent research indicates that investments in player experience (PX)
can result in substantial returns on investment (ROI). By prioritizing PX and harnessing
the capabilities of AI...
DeepInfra is now a supported Hugging Face Inference ProviderDeepInfra is officially live as an Inference Provider on the Hugging Face Hub. You can now call DeepInfra-hosted models directly from Hugging Face model pages, through our OpenAI-compatible router (use it with any OpenAI SDK), or via the Hugging Face SDKs in Python and JavaScript.© 2026 DeepInfra. All rights reserved.