DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Qwen logo

Qwen/

Qwen3.8-Flash

Partner

$0.113

in

$0.382

out

$0.014

cached

/ 1M tokens

Qwen's fast, low-cost model with a one-million-token context window, billed at a flat rate across the entire window.

Public
1,000,000
JSON
Function
Qwen/Qwen3.8-Flash cover image
api
Qwen/Qwen3.8-Flash cover image
Qwen3.8 Flash

Ask me anything

0.00s

You need to log in to use this model

Log In

Settings

Model Information
  • The fast, low-cost member of the Qwen3.8 family, built for high-volume work where latency and price matter more than peak capability.
  • One-million-token context window billed at a single flat rate across the whole window, with no long-context price tiers to plan around.
  • Handles tool calling and structured output over an OpenAI-compatible chat completions API.
  • Context caching brings repeated prefixes down to a fraction of the standard input rate, which suits agents and long running conversations.