GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

GLM-5.2

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Kimi-K2.7-Code

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

NVIDIA-Nemotron-3-Ultra-550B-A55B

DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.

DeepSeek-V4-Flash

DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks.

DeepSeek-V4-Pro

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Kimi-K2.6

MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows. 

MiMo-V2.5

MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).

MiMo-V2.5-Pro

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

Qwen3.6-35B-A3B

GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).

GLM-5.1

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

Qwen3.5-397B-A17B

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

gemma-4-26B-A4B-it

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

gemma-4-31B-it

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

NVIDIA-Nemotron-3-Super-120B-A12B

GLM-5 is an advanced, open-source large language model designed for developers tackling the toughest challenges. It excels at long-context reasoning, multi-step tool orchestration, and complex systems engineering, making it the ideal choice for powering sophisticated agents and applications that require high-level cognitive tasks.

GLM-5

  Qwen3-TTS is an advanced text-to-speech model by Alibaba's Qwen team, delivering stable, expressive, and low-latency speech generation across 10 languages.                                                                                                                                                                                                                                                                                                                                           Key capabilities:                                                                                                                                                                                                                                  - 9 preset voices — Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna, Sohee — covering diverse genders, ages, and accents                                                                                                              - Voice cloning — clone any voice from a short (~3s) audio sample via the voice_id parameter   - Instruction control — adjust tone, emotion, and speaking style with natural language (e.g. "speak slowly and calmly", "excited tone")   - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese   - Streaming support — real-time PCM streaming with ~97ms first-byte latency   - Multiple output formats — WAV, MP3, FLAC, PCM    Built on a 1.7B parameter architecture using discrete multi-codebook language modeling for end-to-end speech synthesis without cascading errors. Uses a custom 12Hz acoustic tokenizer that preserves paralinguistic information and   environmental audio details.

Qwen3-TTS

● Qwen3-TTS-VoiceDesign is a voice design variant of Qwen3-TTS by Alibaba's Qwen team. Instead of selecting from preset voices, you describe the voice you want in natural language — and the model generates speech in that voice.                                                                                                                                                                                                                                                                     Key capabilities:                                                                                                                                                                                                                                  - Natural language voice control — describe any voice with free text (e.g. "a deep male voice with a calm, authoritative presence", "a young cheerful female with a warm and friendly tone")   - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese                                                                                                                                         - Streaming support — real-time PCM streaming   - Multiple output formats — WAV, MP3, FLAC, PCM    Built on the same 1.7B parameter architecture as Qwen3-TTS, using discrete multi-codebook language modeling and a custom 12Hz acoustic tokenizer for high-quality end-to-end speech synthesis.

Qwen3-TTS-VoiceDesign

The latest flagship model in the Qwen family. State-of-the-art results across a comprehensive suite of benchmarks — including knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding.

Qwen3-Max

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Qwen3-Max-Thinking

Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.

Kimi-K2.5

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

GLM-4.7-Flash

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.

DeepSeek-V3.2

The fastest model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

FLUX-2-klein-4b

The best quality-to-latency ratio, production apps model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

FLUX-2-klein-9b

claude

Claude

deepseek

DeepSeek

flux

Flux

gemini

Gemini

llama

Llama

mistral

Mistral

nemotron

Nemotron

qwen

Qwen

ACE-Step v1.5 is a powerful open-source music foundation model that turns a text prompt into a complete song — vocals, lyrics, and instrumentation — at quality that rivals commercial tools. We run the high-quality XL checkpoint with its planning step  ("thinking") on by default, so generations favor musical structure and coherence over raw speed.

You can use cURL or any other http client to run inferences:

```bash
curl -X POST \
    -d '{"prompt": "a gentle, melodic acoustic ballad \u2014 soft fingerpicked guitar and warm piano, tender female vocals, calm and hopeful, welcoming the sunrise, around 70 bpm, C major", "lyrics": "[verse]\nThe moon clocks out, the stars go to bed\n[chorus]\nUp pops the sun, all gold and red", "response_format": "mp3", "negative_prompt": null}'  \
    -H "Authorization: bearer $DEEPINFRA_TOKEN"  \
    -H 'Content-Type: application/json'  \
    'https://api.deepinfra.com/v1/inference/ACE-Step/acestep-v15-xl-sft'
```

which will give you back something similar to:

```json
{
  "audio": null,
  "output_format": "mp3",
  "duration_seconds": 0,
  "seed": 0,
  "generated_lyrics": null,
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0,
    "output_length": 0
  }
}

```


You can use our command-line tool [deepctl](/docs/advanced/deepctl) to run
inferences:

```bash
deepctl infer \
    -m 'ACE-Step/acestep-v15-xl-sft'  \
    -i 'prompt=a gentle, melodic acoustic ballad — soft fingerpicked guitar and warm piano, tender female vocals, calm and hopeful, welcoming the sunrise, around 70 bpm, C major'  \
    -i 'lyrics=[verse]
The moon clocks out, the stars go to bed
[chorus]
Up pops the sun, all gold and red'  \
    -i response_format=mp3  \
    -i negative_prompt=None
```

which will give you back something similar to:

```json
{
  "audio": null,
  "output_format": "mp3",
  "duration_seconds": 0,
  "seed": 0,
  "generated_lyrics": null,
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0,
    "output_length": 0
  }
}

```


Describe the music — genre, mood, instrumentation, vocals, tempo and key — and what you'd like the song to be about.

prompt

Lyrics to sing, with optional [verse]/[chorus] structure tags. Empty means the model writes its own lyrics (or, for instrumental requests, none).

lyrics

Length of the track in seconds (30-500). Leave empty to let the model choose.

duration

Output audio format. mp3 is small and universally playable; flac and wav are lossless.

response_format

Let the planning model arrange structure and phrasing before generation. Higher quality but slower; disable it for faster results.

think

Generate instrumental music with no vocals, regardless of lyrics.

instrumental

Seed for reproducible output; empty for random.

seed

Target tempo in beats per minute. Empty means the model chooses.

Language for auto-written lyrics. Ignored when you supply lyrics or for instrumental tracks.

language_code

Styles, genres, moods, or instruments to steer the song away from — a compositional nudge, not an audio filter. Ignored when the model writes its own lyrics.

negative_prompt

Classifier-free guidance strength: higher follows the prompt more strictly, lower is more creative/varied. Default 7.

guidance_scale

Generation task. text2music = from text only. cover = restyle a reference track. cover-nofsq = cover without FSQ quantization. repaint = regenerate a region of the input. Every task except text2music requires input_audio.

task

The track to transform — used by the cover / cover-nofsq / repaint tasks (set Task first). Leave empty for plain text2music. This is the audio whose structure/melody is reworked, not the style reference below.

input_audio

Optional style reference: the output borrows the sonic character (timbre, production, vibe) of this clip WITHOUT copying its melody or structure. Works with any task, including text2music. Strength is set by Audio Cover / Reference Strength below (lower ≈ looser).

reference_audio

How strongly the source/reference audio influences the result — governs both cover (the source track) and reference_audio (the style clip). Lower (~0.2) = looser style transfer; higher (→1.0) = closer to the reference.

audio_cover_strength

cover: morph the source toward your prompt/lyrics (an edit), describing the original via flow_edit_source_*.

flow_edit

flow_edit: caption describing the ORIGINAL input audio.

flow_edit_source_caption

flow_edit: lyrics of the ORIGINAL input audio.

flow_edit_source_lyrics

repaint: start of the region to regenerate, seconds.

repaint_start

repaint: end of the region to regenerate, seconds (-1 = to the end).

repaint_end

repaint: how much the regenerated region may diverge.

repaint_mode

repaint (balanced mode): 0=aggressive, 1=conservative.

repaint_strength

Variation amount when re-generating from the same inputs (0 = none).

retake_variance

Seed for the retake variation (used only when retake_variance > 0).

retake_seed

Advanced: precomputed audio semantic codes for code-controlled generation.

audio_codes

The webhook to call when inference is done, by default you will get the output in the response of your inference request

webhook

ACE-Step music generation: text2music (text-only, the default) plus the
audio-conditioned tasks (cover / cover-nofsq / repaint). A request with
task=text2music and no input_audio is plain text-to-music, so this schema is
a strict superset of the original text-only one — backward compatible.

acestep-v15-xl-sft

HTTP/cURL API

Input fields

Input Schema

Output Schema