GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

GLM-5.2

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Kimi-K2.7-Code

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

NVIDIA-Nemotron-3-Ultra-550B-A55B

Nemotron 3 Nano Omni is an open multimodal model built on a hybrid Mixture-of-Experts (MoE) architecture, engineered for high efficiency and strong accuracy across image, video, audio, and text inputs. It powers always-on sub-agents for computer use, document intelligence, and audio-video understanding—replacing fragmented vision, speech, and language pipelines with a single unified inference pass.

Nemotron-3-Nano-Omni-30B-A3B-Reasoning

DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.

DeepSeek-V4-Flash

DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks.

DeepSeek-V4-Pro

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Kimi-K2.6

MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows. 

MiMo-V2.5

MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).

MiMo-V2.5-Pro

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

Qwen3.6-35B-A3B

GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).

GLM-5.1

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

Qwen3.5-397B-A17B

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

gemma-4-26B-A4B-it

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

gemma-4-31B-it

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

NVIDIA-Nemotron-3-Super-120B-A12B

GLM-5 is an advanced, open-source large language model designed for developers tackling the toughest challenges. It excels at long-context reasoning, multi-step tool orchestration, and complex systems engineering, making it the ideal choice for powering sophisticated agents and applications that require high-level cognitive tasks.

GLM-5

  Qwen3-TTS is an advanced text-to-speech model by Alibaba's Qwen team, delivering stable, expressive, and low-latency speech generation across 10 languages.                                                                                                                                                                                                                                                                                                                                           Key capabilities:                                                                                                                                                                                                                                  - 9 preset voices — Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna, Sohee — covering diverse genders, ages, and accents                                                                                                              - Voice cloning — clone any voice from a short (~3s) audio sample via the voice_id parameter   - Instruction control — adjust tone, emotion, and speaking style with natural language (e.g. "speak slowly and calmly", "excited tone")   - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese   - Streaming support — real-time PCM streaming with ~97ms first-byte latency   - Multiple output formats — WAV, MP3, FLAC, PCM    Built on a 1.7B parameter architecture using discrete multi-codebook language modeling for end-to-end speech synthesis without cascading errors. Uses a custom 12Hz acoustic tokenizer that preserves paralinguistic information and   environmental audio details.

Qwen3-TTS

● Qwen3-TTS-VoiceDesign is a voice design variant of Qwen3-TTS by Alibaba's Qwen team. Instead of selecting from preset voices, you describe the voice you want in natural language — and the model generates speech in that voice.                                                                                                                                                                                                                                                                     Key capabilities:                                                                                                                                                                                                                                  - Natural language voice control — describe any voice with free text (e.g. "a deep male voice with a calm, authoritative presence", "a young cheerful female with a warm and friendly tone")   - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese                                                                                                                                         - Streaming support — real-time PCM streaming   - Multiple output formats — WAV, MP3, FLAC, PCM    Built on the same 1.7B parameter architecture as Qwen3-TTS, using discrete multi-codebook language modeling and a custom 12Hz acoustic tokenizer for high-quality end-to-end speech synthesis.

Qwen3-TTS-VoiceDesign

MiniMax M2.5 is SOTA in coding, agentic tool use and search, office work, and a range of other economically valuable tasks, boasting scores of 80.2% in SWE-Bench Verified, 51.3% in Multi-SWE-Bench, and 76.3% in BrowseComp (with context management).

MiniMax-M2.5

The latest flagship model in the Qwen family. State-of-the-art results across a comprehensive suite of benchmarks — including knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding.

Qwen3-Max

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Qwen3-Max-Thinking

Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.

Kimi-K2.5

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

GLM-4.7-Flash

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.

DeepSeek-V3.2

The fastest model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

FLUX-2-klein-4b

The best quality-to-latency ratio, production apps model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

FLUX-2-klein-9b

claude

Claude

deepseek

DeepSeek

flux

Flux

gemini

Gemini

llama

Llama

mistral

Mistral

nemotron

Nemotron

qwen

Qwen

You can use cURL or any other http client to run inferences:

```bash
curl -X POST \
    -d '{"input": "The quick brown fox jumps over the lazy dog", "voice": "A deep male voice with a calm, authoritative tone"}'  \
    -H "Authorization: bearer $DEEPINFRA_TOKEN"  \
    -H 'Content-Type: application/json'  \
    'https://api.deepinfra.com/v1/inference/Qwen/Qwen3-TTS-VoiceDesign'
```

which will give you back something similar to:

```json
{
  "audio": null,
  "input_character_length": 0,
  "output_format": "",
  "words": [
    {
      "end": 1.0,
      "start": 0.0,
      "text": "Hello"
    },
    {
      "end": 5.0,
      "start": 4.0,
      "text": "World"
    }
  ],
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0,
    "output_length": 0
  }
}

```


You can use our command-line tool [deepctl](/docs/advanced/deepctl) to run
inferences:

```bash
deepctl infer \
    -m 'Qwen/Qwen3-TTS-VoiceDesign'  \
    -i 'input=The quick brown fox jumps over the lazy dog'  \
    -i 'voice=A deep male voice with a calm, authoritative tone'
```

which will give you back something similar to:

```json
{
  "audio": null,
  "input_character_length": 0,
  "output_format": "",
  "words": [
    {
      "end": 1.0,
      "start": 0.0,
      "text": "Hello"
    },
    {
      "end": 5.0,
      "start": 4.0,
      "text": "World"
    }
  ],
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0,
    "output_length": 0
  }
}

```


The DeepInfra OpenAI-compatible Speech API endpoint enables users to effortlessly convert text into speech audio. This document outlines how to integrate and utilize this endpoint to quickly create speech from text inputs, leveraging various audio output formats.

## Create speech

Use the following example of pythong code to generate an audio file from your text input:

```python
from pathlib import Path
from openai import OpenAI

client = OpenAI(base_url="https://api.deepinfra.com/v1/openai",
                api_key="$DEEPINFRA_TOKEN")

speech_file_path = Path(__file__).parent / "speech.mp3"
with client.audio.speech.with_streaming_response.create(
  model="Qwen/Qwen3-TTS-VoiceDesign",
  voice="A young female speaker with a clear, warm voice",
  input="The quick brown fox jumped over the lazy dog.",
  response_format="mp3",
) as response:
  response.stream_to_file(speech_file_path)
```

The API returns the generated audio file in the requested format (e.g., `mp3`, `pcm`). The example above saves the audio output directly as `speech.mp3`.

### Supported Audio Formats

| Format | Description                           | Recommended Usage                          |
|--------|---------------------------------------|--------------------------------------------|
| pcm    | Raw, uncompressed audio               | *Best for lowest latency streaming*        |
| mp3    | Compressed audio suitable for storage | General use and file distribution          |
| opus   | Compressed audio ideal for streaming  | Efficient streaming and real-time playback |
| flac   | Lossless audio compression            | High-quality archival storage              |
| wav    | Uncompressed audio                    | High-quality audio processing tasks        |

### Performance Recommendations

- **For lowest latency (fastest initial audio chunk):**
  - Use the `pcm` output format.

- **For general purposes:**
  - Use the `mp3` or `opus` formats for optimal balance between quality and file size.



The DeepInfra OpenAI-compatible Speech API endpoint enables users to effortlessly convert text into speech audio. This document outlines how to integrate and utilize this endpoint to quickly create speech from text inputs, leveraging various audio output formats.

## Create speech

Use the following example of js code to generate an audio file from your text input:

```javascript
import fs from "fs";
import path from "path";
import OpenAI from "openai";

const openai = new OpenAI(base_url="https://api.deepinfra.com/v1/openai",
                          api_key="$DEEPINFRA_TOKEN");

const speechFile = path.resolve("./speech.mp3");

async function main() {
  const mp3 = await openai.audio.speech.create({
    model: "Qwen/Qwen3-TTS-VoiceDesign",
    voice: "A young female speaker with a clear, warm voice",
    input: "The quick brown fox jumped over the lazy dog.",
    response_format: "mp3",
  });
  console.log(speechFile);
  const buffer = Buffer.from(await mp3.arrayBuffer());
  await fs.promises.writeFile(speechFile, buffer);
}
main();
```

The API returns the generated audio file in the requested format (e.g., `mp3`, `pcm`). The example above saves the audio output directly as `speech.mp3`.

### Supported Audio Formats

| Format | Description                           | Recommended Usage                          |
|--------|---------------------------------------|--------------------------------------------|
| pcm    | Raw, uncompressed audio               | *Best for lowest latency streaming*        |
| mp3    | Compressed audio suitable for storage | General use and file distribution          |
| opus   | Compressed audio ideal for streaming  | Efficient streaming and real-time playback |
| flac   | Lossless audio compression            | High-quality archival storage              |
| wav    | Uncompressed audio                    | High-quality audio processing tasks        |

### Performance Recommendations

- **For lowest latency (fastest initial audio chunk):**
  - Use the `pcm` output format.

- **For general purposes:**
  - Use the `mp3` or `opus` formats for optimal balance between quality and file size.



The DeepInfra OpenAI-compatible Speech API endpoint enables users to effortlessly convert text into speech audio. This document outlines how to integrate and utilize this endpoint to quickly create speech from text inputs, leveraging various audio output formats.

## Create speech

Use the following example `curl` request to generate an audio file from your text input:

```bash
curl https://api.deepinfra.com/v1/openai/audio/speech \
  -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-TTS-VoiceDesign",
    "input": "The quick brown fox jumped over the lazy dog.",
    "voice": "A young female speaker with a clear, warm voice",
    "response_format": "mp3"
  }' \
  --output speech.mp3
```

The API returns the generated audio file in the requested format (e.g., `mp3`, `pcm`). The example above saves the audio output directly as `speech.mp3`.

### Supported Audio Formats

| Format | Description                           | Recommended Usage                          |
|--------|---------------------------------------|--------------------------------------------|
| pcm    | Raw, uncompressed audio               | *Best for lowest latency streaming*        |
| mp3    | Compressed audio suitable for storage | General use and file distribution          |
| opus   | Compressed audio ideal for streaming  | Efficient streaming and real-time playback |
| flac   | Lossless audio compression            | High-quality archival storage              |
| wav    | Uncompressed audio                    | High-quality audio processing tasks        |

### Performance Recommendations

- **For lowest latency (fastest initial audio chunk):**
  - Use the `pcm` output format.

- **For general purposes:**
  - Use the `mp3` or `opus` formats for optimal balance between quality and file size.



The DeepInfra ElevenLabs-compatible Speech API endpoint allows users to seamlessly convert text inputs into high-quality speech audio. This document provides guidance on integrating and using this endpoint to generate realistic audio files efficiently.

Use the following examples of `py` request to generate an audio file from your text input:

## Create Speech (non-streaming)

```bash
from elevenlabs import ElevenLabs

client = ElevenLabs(
    api_key="$DEEPINFRA_TOKEN",
    base_url="https://api.deepinfra.com/",
)
client.text_to_speech.convert(
    voice_id="A young female speaker with a clear, warm voice",
    output_format="mp3",
    text="The quick brown fox jumped over the lazy dog.",
    model_id="Qwen/Qwen3-TTS-VoiceDesign",
)
```

## Create Speech with Streaming

```bash
from elevenlabs import ElevenLabs

client = ElevenLabs(
    api_key="$DEEPINFRA_TOKEN",
)
client.text_to_speech.convert_as_stream(
    voice_id="A young female speaker with a clear, warm voice",
    output_format="pcm",
    text="The quick brown fox jumped over the lazy dog.",
    model_id="Qwen/Qwen3-TTS-VoiceDesign",
)
```

The API returns the generated audio in the requested format, such as `mp3` or `pcm`. The example above saves the audio directly as `speech.mp3`.

### Supported Audio Formats

| Format | Description                           | Recommended Usage                          |
|--------|---------------------------------------|--------------------------------------------|
| pcm    | Raw, uncompressed audio               | *Lowest latency streaming scenarios*       |
| mp3    | Compressed audio suitable for storage | General use and easy file sharing          |
| opus   | Compressed audio ideal for streaming  | Real-time audio streaming applications     |
| flac   | Lossless audio compression            | Archival storage and high-fidelity needs   |
| wav    | Uncompressed audio                    | Audio processing and editing tasks         |

### Performance Recommendations

- **Lowest latency:** Using the `pcm` output format for streaming applications is **HIGHLY RECOMMENDED**.
- **General use:** Use `mp3` or `opus` formats for the best trade-off between quality and file size.



The DeepInfra ElevenLabs-compatible Speech API endpoint allows users to seamlessly convert text inputs into high-quality speech audio. This document provides guidance on integrating and using this endpoint to generate realistic audio files efficiently.

Use the following examples of `js` request to generate an audio file from your text input:

## Create Speech (non-streaming)

```bash
import { ElevenLabsClient } from "elevenlabs";

const client = new ElevenLabsClient({ apiKey: "$DEEPINFRA_TOKEN", base_url: "https://api.deepinfra.com/" });
await client.textToSpeech.convert("A young female speaker with a clear, warm voice", {
    output_format: "mp3",
    text: "The quick brown fox jumped over the lazy dog.",
    model_id: "Qwen/Qwen3-TTS-VoiceDesign"
});
```

## Create Speech with Streaming

```bash
import { ElevenLabsClient } from "elevenlabs";

const client = new ElevenLabsClient({ apiKey: "$DEEPINFRA_TOKEN" });
await client.textToSpeech.convert("A young female speaker with a clear, warm voice", {
    output_format: "pcm",
    text: "The quick brown fox jumped over the lazy dog.",
    model_id: "Qwen/Qwen3-TTS-VoiceDesign"
});
```

The API returns the generated audio in the requested format, such as `mp3` or `pcm`. The example above saves the audio directly as `speech.mp3`.

### Supported Audio Formats

| Format | Description                           | Recommended Usage                          |
|--------|---------------------------------------|--------------------------------------------|
| pcm    | Raw, uncompressed audio               | *Lowest latency streaming scenarios*       |
| mp3    | Compressed audio suitable for storage | General use and easy file sharing          |
| opus   | Compressed audio ideal for streaming  | Real-time audio streaming applications     |
| flac   | Lossless audio compression            | Archival storage and high-fidelity needs   |
| wav    | Uncompressed audio                    | Audio processing and editing tasks         |

### Performance Recommendations

- **Lowest latency:** Using the `pcm` output format for streaming applications is **HIGHLY RECOMMENDED**.
- **General use:** Use `mp3` or `opus` formats for the best trade-off between quality and file size.



The DeepInfra ElevenLabs-compatible Speech API endpoint allows users to seamlessly convert text inputs into high-quality speech audio. This document provides guidance on integrating and using this endpoint to generate realistic audio files efficiently.

Use the following examples of `curl` request to generate an audio file from your text input:

## Create Speech (non-streaming)

```bash
curl -X POST "https://api.deepinfra.com/v1/text-to-speech/A young female speaker with a clear, warm voice" \
     -H "xi-api-key: $DEEPINFRA_TOKEN" \
     -H "Content-Type: application/json" \
     -d '{
  "text": "The quick brown fox jumped over the lazy dog.",
  "model_id": "Qwen/Qwen3-TTS-VoiceDesign",
  "output_format": "mp3",
}' --output speech.mp3
```

## Create Speech with Streaming

```bash
curl -X POST "https://api.deepinfra.com/v1/text-to-speech/A young female speaker with a clear, warm voice/stream" \
     -H "xi-api-key: $DEEPINFRA_TOKEN" \
     -H "Content-Type: application/json" \
     -d '{
  "text": "The quick brown fox jumped over the lazy dog.",
  "model_id": "Qwen/Qwen3-TTS-VoiceDesign",
  "output_format": "pcm",
}' --output speech.pcm
```

The API returns the generated audio in the requested format, such as `mp3` or `pcm`. The example above saves the audio directly as `speech.mp3`.

### Supported Audio Formats

| Format | Description                           | Recommended Usage                          |
|--------|---------------------------------------|--------------------------------------------|
| pcm    | Raw, uncompressed audio               | *Lowest latency streaming scenarios*       |
| mp3    | Compressed audio suitable for storage | General use and easy file sharing          |
| opus   | Compressed audio ideal for streaming  | Real-time audio streaming applications     |
| flac   | Lossless audio compression            | Archival storage and high-fidelity needs   |
| wav    | Uncompressed audio                    | Audio processing and editing tasks         |

### Performance Recommendations

- **Lowest latency:** Using the `pcm` output format for streaming applications is **HIGHLY RECOMMENDED**.
- **General use:** Use `mp3` or `opus` formats for the best trade-off between quality and file size.



The service tier used for processing the request. 'priority' processes the request with higher priority (premium rate); 'flex' processes it at lower priority for a discount, served only when spare capacity exists and may be retried/timed out under load. Both apply only to models that support the respective tier.

service_tier

input

Natural language description of the desired voice (e.g. "A young cheerful female with a warm tone")

voice

Language of the speech. Use Auto for auto-detection.

language

response_format

The webhook to call when inference is done, by default you will get the output in the response of your inference request

webhook

Select the desired language for the speech output. Use Auto for auto-detection.

Qwen3TtsLanguage

ServiceTier

Select the desired format for the speech output. Supported formats include mp3, opus, flac, wav, and pcm.

Qwen3-TTS-VoiceDesign

HTTP/cURL API

Input fields

Input Schema

Output Schema