DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Qwen3.8-27B: Frontier-Level Vision-Language AI

Published on 2026.10.09 by DeepInfra
Qwen3.8-27B: Frontier-Level Vision-Language AI

Qwen3.8-27B is a 27-billion parameter dense vision-language model built to close the gap between compact deployment and frontier-level performance. It is engineered for complex, multi-step tasks, including coding, professional research, and long-horizon agentic workflows.

The model pairs a native vision encoder with a 262,144-token context window, extensible to 1,000,000 tokens. It reads both text and visual data, from STEM diagrams to hour-scale video. It suits teams building autonomous agents or document intelligence tools that reason over images and long documents in the same request.

Latest Developments and Architecture

Qwen3.8-27B is built on the architectural foundation of the Qwen3.5 series and introduces several changes over its predecessor, Qwen3.6-27B.

Under the hood, it is a 64-layer causal language model with a vision encoder attached, with a hidden dimension of 5,120. The layout mixes linear-attention (Gated DeltaNet) blocks with full-attention (Gated Attention) blocks, and multi-token prediction is trained in to speed up generation.

Flexible thinking control

One of the most significant additions in Qwen3.8 is flexible thinking control, which gives developers granular command over the model’s reasoning process. Thinking mode is on by default and can be disabled per request, or tuned rather than switched off entirely:

  • Reasoning depth: Use the reasoning_effort parameter to align compute resources with task complexity.
  • Thinking persistence: The preserve_thinking parameter retains reasoning context from earlier messages, improving continuity in long-running agentic workflows.

Native multimodal intelligence

Rather than relying on external vision plugins, Qwen3.8-27B has native support for image and video understanding. It processes complex documents and visual data within its 262,144-token native context window.

Performance Benchmarks and Comparative Analysis

Qwen3.8-27B outperforms Qwen3.6-27B and Muse Glimmer-30B across most text and multimodal evaluations, and beats the larger Opus4.6 Max on several agentic and coding benchmarks.

Text and reasoning performance

BenchmarkQwen3.8-27BQwen3.6-27BMuse Glimmer-30BOpus4.6 Max
GPQA Diamond (scientific reasoning)89.287.883.591.3
SWE-bench Pro (agentic coding)61.753.551.253.4
LiveCodeBench v6 (competitive coding)90.383.9N/A88.8
IFBench (instruction following)79.569.177.062.5
CoWorkBench (long-horizon office work)70.761.0N/A68.2

Vision-language performance

BenchmarkQwen3.8-27BQwen3.6-27BMuse Glimmer-30BOpus4.6 Max
OSWorld-Verified (computer use)84.363.965.972.7
MathVision, with reasoning (visual math)94.685.1N/A65.5
OmniDocBench 1.5 (document intelligence)91.189.4N/A86.6
RealWorldQA (perception)85.984.1N/A73.9

Key highlights

  • Agentic execution: The model leads OSWorld-Verified for UI navigation at 84.3. It also ties the strongest result on multimodal tool use, with a Pass@3 of 57.4 on ClawEval-MM.
  • Software engineering: A score of 79.0 on the internal QwenSWEBench is the best in its comparison set, well ahead of Qwen3.6-27B at 49.3 and Opus4.6 Max at 63.8. It marks a substantial jump in handling complex, multi-file codebases.
  • Efficiency: Multi-token prediction is trained into the architecture, which typically improves inference speed by generating more than one token per forward pass.

Getting Started with the API

DeepInfra provides an OpenAI-compatible API to integrate Qwen3.8-27B into your applications.

Authentication

To interact with the API, use your DeepInfra API key in the HTTP headers:

Authorization: Bearer <YOUR_DEEPINFRA_API_KEY>
Content-Type: application/json
copy

API endpoint basics

  • Endpoint: https://api.deepinfra.com/v1/openai/chat/completions
  • Model identifier: Qwen/Qwen3.8-27B

The full parameter reference, including streaming, is on the Qwen3.8-27B API page.

Using cURL

curl https://api.deepinfra.com/v1/openai/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $DEEPINFRA_API_KEY" \
  -d '{
    "model": "Qwen/Qwen3.8-27B",
    "messages": [
      {
        "role": "user",
        "content": "Write a Python function to check if a number is prime."
      }
    ]
  }'
copy

Using Python

import requests
import os

API_KEY = os.getenv("DEEPINFRA_API_KEY")
url = "https://api.deepinfra.com/v1/openai/chat/completions"

headers = {
    "Authorization": f"Bearer {API_KEY}",
    "Content-Type": "application/json"
}

payload = {
    "model": "Qwen/Qwen3.8-27B",
    "messages": [{"role": "user", "content": "Explain the benefits of the Qwen3.8 architecture."}]
}

response = requests.post(url, headers=headers, json=payload)
print(response.json())
copy

Advanced Configuration Parameters

Customize the model’s behavior with the following parameters in the JSON payload:

ParameterTypeDescription
modelstringRequired. Set to Qwen/Qwen3.8-27B.
messagesarrayRequired. Supports text, image, and video inputs.
reasoning_effortstring / floatTunes reasoning depth when thinking mode is active.
preserve_thinkingbooleanRetains reasoning context across multi-turn conversations.
response_formatobjectUse {“type”: “json_object”} for structured JSON output.
streambooleanSet to true to receive the response as a token stream.

Pricing and Cost Optimization

DeepInfra prices Qwen3.8-27B per token, with a 25% promotional discount currently applied to the standard tier and a further break on cached input. Cached input already runs 75% below the standard input rate at list price, which matters for repetitive prompts such as reused system messages or tool schemas.

TierInputOutputCached Input
Standard (list)$0.20$2.50$0.05
Standard (current, 25% off)$0.15$1.875$0.038
Priority (1.5x)$0.225$2.8125$0.0563
Flex (0.8x)$0.12$1.50$0.03

Priority runs at 1.5x the standard rate for lower latency, and Flex runs at 0.8x for workloads that can tolerate slower response times. All tiers are shown per 1 million tokens. For volume discounts or a private endpoint, check the pricing page or your DeepInfra dashboard.

Conclusion

Qwen3.8-27B combines a 1M-token extensible context window with flexible thinking control and strong multimodal reasoning in a 27B dense model. It is a capable option for demanding agentic and document-heavy workloads. Teams that need more headroom can also look at the larger Qwen3.8-2.4T-A95B, the sparse mixture-of-experts variant of Qwen3.8-Max.

Next steps

  • Explore multimodal inputs: Pass images or video frames within the messages array for visual analysis.
  • Deploy privately: For dedicated resources and an isolated environment, deploy a private endpoint.
  • Stay updated: Visit the DeepInfra documentation for the latest SDKs and integration guides.
Related articles
We Benchmarked NVIDIA Vera, the CPU for Agents. Here's What We MeasuredWe Benchmarked NVIDIA Vera, the CPU for Agents. Here's What We MeasuredDeepInfra runs AI agents in production, so when NVIDIA built a CPU for agents, we measured it ourselves with our own harness, our own agent, and a methodology we locked before the hardware arrived.
Best Kimi K3 SaaS Tools & API PlatformsBest Kimi K3 SaaS Tools & API Platforms<p>Kimi K3, with its 2.8 trillion parameters and 1M-token context window, represents a significant leap in large language model capabilities. However, deploying, accessing, and managing a model of this scale presents real infrastructure challenges. From optimizing inference latency and managing GPU compute costs to handling multimodal vision capabilities, selecting the right deployment platform is critical [&hellip;]</p>
DeepSeek-V4.1-Flash Is Now on DeepInfraDeepSeek-V4.1-Flash Is Now on DeepInfra<p>DeepSeek’s new V4.1-Flash uses just 8 billion active parameters during input processing, out of a 552-billion-parameter backbone, and still beats the much larger DeepSeek-V4-Pro across every agentic benchmark DeepSeek published. Released in September 2026, it is built around a Causal Encoder-Decoder design that makes the gap between total and active parameters possible. For developers running [&hellip;]</p>