DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Qwen3.8-27B is a 27-billion parameter dense vision-language model built to close the gap between compact deployment and frontier-level performance. It is engineered for complex, multi-step tasks, including coding, professional research, and long-horizon agentic workflows.
The model pairs a native vision encoder with a 262,144-token context window, extensible to 1,000,000 tokens. It reads both text and visual data, from STEM diagrams to hour-scale video. It suits teams building autonomous agents or document intelligence tools that reason over images and long documents in the same request.
Qwen3.8-27B is built on the architectural foundation of the Qwen3.5 series and introduces several changes over its predecessor, Qwen3.6-27B.
Under the hood, it is a 64-layer causal language model with a vision encoder attached, with a hidden dimension of 5,120. The layout mixes linear-attention (Gated DeltaNet) blocks with full-attention (Gated Attention) blocks, and multi-token prediction is trained in to speed up generation.
One of the most significant additions in Qwen3.8 is flexible thinking control, which gives developers granular command over the model’s reasoning process. Thinking mode is on by default and can be disabled per request, or tuned rather than switched off entirely:
Rather than relying on external vision plugins, Qwen3.8-27B has native support for image and video understanding. It processes complex documents and visual data within its 262,144-token native context window.
Qwen3.8-27B outperforms Qwen3.6-27B and Muse Glimmer-30B across most text and multimodal evaluations, and beats the larger Opus4.6 Max on several agentic and coding benchmarks.
| Benchmark | Qwen3.8-27B | Qwen3.6-27B | Muse Glimmer-30B | Opus4.6 Max |
|---|---|---|---|---|
| GPQA Diamond (scientific reasoning) | 89.2 | 87.8 | 83.5 | 91.3 |
| SWE-bench Pro (agentic coding) | 61.7 | 53.5 | 51.2 | 53.4 |
| LiveCodeBench v6 (competitive coding) | 90.3 | 83.9 | N/A | 88.8 |
| IFBench (instruction following) | 79.5 | 69.1 | 77.0 | 62.5 |
| CoWorkBench (long-horizon office work) | 70.7 | 61.0 | N/A | 68.2 |
| Benchmark | Qwen3.8-27B | Qwen3.6-27B | Muse Glimmer-30B | Opus4.6 Max |
|---|---|---|---|---|
| OSWorld-Verified (computer use) | 84.3 | 63.9 | 65.9 | 72.7 |
| MathVision, with reasoning (visual math) | 94.6 | 85.1 | N/A | 65.5 |
| OmniDocBench 1.5 (document intelligence) | 91.1 | 89.4 | N/A | 86.6 |
| RealWorldQA (perception) | 85.9 | 84.1 | N/A | 73.9 |
DeepInfra provides an OpenAI-compatible API to integrate Qwen3.8-27B into your applications.
To interact with the API, use your DeepInfra API key in the HTTP headers:
Authorization: Bearer <YOUR_DEEPINFRA_API_KEY>
Content-Type: application/jsonThe full parameter reference, including streaming, is on the Qwen3.8-27B API page.
curl https://api.deepinfra.com/v1/openai/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $DEEPINFRA_API_KEY" \
-d '{
"model": "Qwen/Qwen3.8-27B",
"messages": [
{
"role": "user",
"content": "Write a Python function to check if a number is prime."
}
]
}'Using Python
import requests
import os
API_KEY = os.getenv("DEEPINFRA_API_KEY")
url = "https://api.deepinfra.com/v1/openai/chat/completions"
headers = {
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json"
}
payload = {
"model": "Qwen/Qwen3.8-27B",
"messages": [{"role": "user", "content": "Explain the benefits of the Qwen3.8 architecture."}]
}
response = requests.post(url, headers=headers, json=payload)
print(response.json())Customize the model’s behavior with the following parameters in the JSON payload:
| Parameter | Type | Description |
|---|---|---|
| model | string | Required. Set to Qwen/Qwen3.8-27B. |
| messages | array | Required. Supports text, image, and video inputs. |
| reasoning_effort | string / float | Tunes reasoning depth when thinking mode is active. |
| preserve_thinking | boolean | Retains reasoning context across multi-turn conversations. |
| response_format | object | Use {“type”: “json_object”} for structured JSON output. |
| stream | boolean | Set to true to receive the response as a token stream. |
DeepInfra prices Qwen3.8-27B per token, with a 25% promotional discount currently applied to the standard tier and a further break on cached input. Cached input already runs 75% below the standard input rate at list price, which matters for repetitive prompts such as reused system messages or tool schemas.
| Tier | Input | Output | Cached Input |
|---|---|---|---|
| Standard (list) | $0.20 | $2.50 | $0.05 |
| Standard (current, 25% off) | $0.15 | $1.875 | $0.038 |
| Priority (1.5x) | $0.225 | $2.8125 | $0.0563 |
| Flex (0.8x) | $0.12 | $1.50 | $0.03 |
Priority runs at 1.5x the standard rate for lower latency, and Flex runs at 0.8x for workloads that can tolerate slower response times. All tiers are shown per 1 million tokens. For volume discounts or a private endpoint, check the pricing page or your DeepInfra dashboard.
Qwen3.8-27B combines a 1M-token extensible context window with flexible thinking control and strong multimodal reasoning in a 27B dense model. It is a capable option for demanding agentic and document-heavy workloads. Teams that need more headroom can also look at the larger Qwen3.8-2.4T-A95B, the sparse mixture-of-experts variant of Qwen3.8-Max.
We Benchmarked NVIDIA Vera, the CPU for Agents. Here's What We MeasuredDeepInfra runs AI agents in production, so when NVIDIA built a CPU for agents, we measured it ourselves with our own harness, our own agent, and a methodology we locked before the hardware arrived.
Best Kimi K3 SaaS Tools & API Platforms<p>Kimi K3, with its 2.8 trillion parameters and 1M-token context window, represents a significant leap in large language model capabilities. However, deploying, accessing, and managing a model of this scale presents real infrastructure challenges. From optimizing inference latency and managing GPU compute costs to handling multimodal vision capabilities, selecting the right deployment platform is critical […]</p>
DeepSeek-V4.1-Flash Is Now on DeepInfra<p>DeepSeek’s new V4.1-Flash uses just 8 billion active parameters during input processing, out of a 552-billion-parameter backbone, and still beats the much larger DeepSeek-V4-Pro across every agentic benchmark DeepSeek published. Released in September 2026, it is built around a Causal Encoder-Decoder design that makes the gap between total and active parameters possible. For developers running […]</p>
© 2026 DeepInfra. All rights reserved.