GLM-5.2 is Z-AI's latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a **solid 1M-token context**.

GLM-5.2

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Kimi-K2.7-Code

Nemotron 3 Ultra is built for, frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

NVIDIA-Nemotron-3-Ultra-550B-A55B

DeepSeek V4 Flash is an efficiency-focused MoE model with 284B total parameters (13B active) and a 1M-token context window. It's tuned for fast inference and high-throughput use cases while still holding up on reasoning and coding tasks.

DeepSeek-V4-Flash

DeepSeek V4 Pro is an MoE model with 1.6T total parameters (49B active) and a 1M-token context window. It's built for advanced reasoning, coding, and long-running agent tasks, and performs well on knowledge, math, and software engineering benchmarks.

DeepSeek-V4-Pro

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

Kimi-K2.6

MiMo-V2.5 is a native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built upon the MiMo-V2-Flash backbone and extended with dedicated vision and audio encoders, it delivers robust performance across multimodal perception, long-context reasoning, and agentic workflows. 

MiMo-V2.5

MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layers Multi-Token Prediction (MTP) introduced in [MiMo-V2-Flash](https://github.com/XiaomiMiMo/MiMo-V2-Flash).

MiMo-V2.5-Pro

Qwen3.6-35B-A3B is Alibaba's latest flagship Mixture-of-Experts model, with 35B total parameters and only 3B activated per token (256 experts, 8 routed + 1 shared). Built on direct feedback from the community, Qwen3.6 prioritizes stability and real-world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience.

Qwen3.6-35B-A3B

GLM-5.1 is Z-AI's next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).

GLM-5.1

Qwen3.5-397B-A17B is Alibaba's most capable Qwen3.5 model, a Mixture-of-Experts architecture with 397B total parameters and 17B activated per token. It features a 262K token context window (extensible to 1M with YaRN), thinking/reasoning mode, tool calling with MCP integration, and support for 201 languages. Sets state-of-the-art results on reasoning, coding, math, and multimodal benchmarks.

Qwen3.5-397B-A17B

Efficient, MoE variant of Gemma 4. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

gemma-4-26B-A4B-it

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input and generating text output.

gemma-4-31B-it

NVIDIA Nemotron 3 Super is a hybrid Mixture-of-Experts (MoE) model engineered for highest compute efficiency and accuracy in multi-agent applications and specialized agentic systems. It is optimized to run many collaborating agents per application on a single GPU, delivering high accuracy for reasoning, tool use, and instruction following.

NVIDIA-Nemotron-3-Super-120B-A12B

GLM-5 is an advanced, open-source large language model designed for developers tackling the toughest challenges. It excels at long-context reasoning, multi-step tool orchestration, and complex systems engineering, making it the ideal choice for powering sophisticated agents and applications that require high-level cognitive tasks.

GLM-5

  Qwen3-TTS is an advanced text-to-speech model by Alibaba's Qwen team, delivering stable, expressive, and low-latency speech generation across 10 languages.                                                                                                                                                                                                                                                                                                                                           Key capabilities:                                                                                                                                                                                                                                  - 9 preset voices — Vivian, Serena, Uncle_Fu, Dylan, Eric, Ryan, Aiden, Ono_Anna, Sohee — covering diverse genders, ages, and accents                                                                                                              - Voice cloning — clone any voice from a short (~3s) audio sample via the voice_id parameter   - Instruction control — adjust tone, emotion, and speaking style with natural language (e.g. "speak slowly and calmly", "excited tone")   - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese   - Streaming support — real-time PCM streaming with ~97ms first-byte latency   - Multiple output formats — WAV, MP3, FLAC, PCM    Built on a 1.7B parameter architecture using discrete multi-codebook language modeling for end-to-end speech synthesis without cascading errors. Uses a custom 12Hz acoustic tokenizer that preserves paralinguistic information and   environmental audio details.

Qwen3-TTS

● Qwen3-TTS-VoiceDesign is a voice design variant of Qwen3-TTS by Alibaba's Qwen team. Instead of selecting from preset voices, you describe the voice you want in natural language — and the model generates speech in that voice.                                                                                                                                                                                                                                                                     Key capabilities:                                                                                                                                                                                                                                  - Natural language voice control — describe any voice with free text (e.g. "a deep male voice with a calm, authoritative presence", "a young cheerful female with a warm and friendly tone")   - 10 languages — English, Chinese, Japanese, Korean, German, French, Russian, Spanish, Italian, Portuguese                                                                                                                                         - Streaming support — real-time PCM streaming   - Multiple output formats — WAV, MP3, FLAC, PCM    Built on the same 1.7B parameter architecture as Qwen3-TTS, using discrete multi-codebook language modeling and a custom 12Hz acoustic tokenizer for high-quality end-to-end speech synthesis.

Qwen3-TTS-VoiceDesign

The latest flagship model in the Qwen family. State-of-the-art results across a comprehensive suite of benchmarks — including knowledge, reasoning, coding, instruction following, human preference alignment, agent tasks, and multilingual understanding.

Qwen3-Max

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

Qwen3-Max-Thinking

Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.

Kimi-K2.5

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

GLM-4.7-Flash

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.

DeepSeek-V3.2

The fastest model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

FLUX-2-klein-4b

The best quality-to-latency ratio, production apps model of the Flux 2 family. Frontier visual intelligence — state-of-the-art image generation and editing from Black Forest Labs

FLUX-2-klein-9b

claude

Claude

deepseek

DeepSeek

flux

Flux

gemini

Gemini

llama

Llama

mistral

Mistral

nemotron

Nemotron

qwen

Qwen

You can POST to our OpenAI Chat Completions compatible endpoint.

#### Simple messages and prompts

Given a list of messages from a conversation, the model will return a response.

```bash
curl "https://api.deepinfra.com/v1/openai/chat/completions" \
 -H "Content-Type: application/json" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -d '{
 "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 "messages": [
 {
 "role": "user",
 "content": "Hello!"
 }
 ]
 }'
```

To which you'd get something like:

```json
{
 "id": "chatcmpl-guMTxWgpFf",
 "object": "chat.completion",
 "created": 1694623155,
 "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 "choices": [
 {
 "index": 0,
 "message": {
 "role": "assistant",
 "content": " Hello! It's nice to meet you. Is there something I can help you with or would you like to chat for a bit?"
 },
 "finish_reason": "stop"
 }
 ],
 "usage": {
 "prompt_tokens": 15,
 "completion_tokens": 16,
 "total_tokens": 31,
 "estimated_cost": 0.0000268
 }
}
```

#### Conversations

To create a longer chat-like conversation you just have to add each response message and each of the user messages to every request. This way the model will have the context and will be able to provide better answers. You can tweak it even further by providing a system message.

```bash
curl "https://api.deepinfra.com/v1/openai/chat/completions" \
 -H "Content-Type: application/json" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -d '{
 "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 "messages": [
 {
 "role": "system",
 "content": "Respond like a michelin starred chef."
 },
 {
 "role": "user",
 "content": "Can you name at least two different techniques to cook lamb?"
 },
 {
 "role": "assistant",
 "content": "Bonjour! Let me tell you, my friend, cooking lamb is an art form, and I'"'"'m more than happy to share with you not two, but three of my favorite techniques to coax out the rich, unctuous flavors and tender textures of this majestic protein. First, we have the classic \"Sous Vide\" method. Next, we have the ancient art of \"Sous le Sable\". And finally, we have the more modern technique of \"Hot Smoking.\""
 },
 {
 "role": "user",
 "content": "Tell me more about the second method."
 }
 ]
 }'
```

The conversation above might return something like the following

```json
{
 "id": "chatcmpl-b23a3fb60cde42ce8f24bb980b4dee87",
 "object": "chat.completion",
 "created": 1715688169,
 "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 "choices": [
 {
 "index": 0,
 "message": {
 "role": "assistant",
 "content": "Sous le Sable, my friend! It's an ancient technique that's been used for centuries in the Middle East and North Africa. The name itself..."
 },
 "finish_reason": "stop"
 }
 ],
 "usage": {
 "prompt_tokens": 149,
 "total_tokens": 487,
 "completion_tokens": 338,
 "estimated_cost": 0.00035493
 }
}
```

The longer the conversation gets, the more time it takes the model to generate the response. The number of messages that you can have in a conversation is limited by the context size of a model. Larger models also usually take more time to respond.

 

### Streaming

You can turn any of the requests above into a streaming request by passing `"stream": true`:

```bash
curl "https://api.deepinfra.com/v1/openai/chat/completions" \
 -H "Content-Type: application/json" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -d '{
 "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 "stream": true,
 "messages": [
 {
 "role": "user",
 "content": "Hello!"
 }
 ]
 }'
```

to which you'd get a sequence of [SSE](https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events) events, finishing with `[DONE]`.

```
data: {"id": "Rc5hsIPHOSfMP3rNSFUw9tfR", "object": "chat.completion.chunk", "created": 1694623354, "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B", "choices": [{"index": 0, "delta": {"role": "assistant", "content": " "}, "finish_reason": null}]}

data: {"id": "Rc5hsIPHOSfMP3rNSFUw9tfR", "object": "chat.completion.chunk", "created": 1694623354, "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B", "choices": [{"index": 0, "delta": {"role": "assistant", "content": " Hi"}, "finish_reason": null}]}

data: {"id": "Rc5hsIPHOSfMP3rNSFUw9tfR", "object": "chat.completion.chunk", "created": 1694623354, "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B", "choices": [{"index": 0, "delta": {"role": "assistant", "content": "!"}, "finish_reason": null}]}

data: {"id": "Rc5hsIPHOSfMP3rNSFUw9tfR", "object": "chat.completion.chunk", "created": 1694623354, "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B", "choices": [{"index": 0, "delta": {"role": "assistant", "content": ""}, "finish_reason": null}]}

data: {"id": "Rc5hsIPHOSfMP3rNSFUw9tfR", "object": "chat.completion.chunk", "created": 1694623354, "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B", "choices": [{"index": 0, "delta": {"role": "assistant", "content": "</s>"}, "finish_reason": null}]}

data: {"id": "Rc5hsIPHOSfMP3rNSFUw9tfR", "object": "chat.completion.chunk", "created": 1694623354, "model": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B", "choices": [{"index": 0, "delta": {}, "finish_reason": "stop"}]}

data: [DONE]
```

You can use the official openai python client to run inferences with us

```python
# Assume openai>=1.0.0
from openai import OpenAI

# Create an OpenAI client with your deepinfra token and endpoint
openai = OpenAI(
 api_key="$DEEPINFRA_TOKEN",
 base_url="https://api.deepinfra.com/v1/openai",
)

chat_completion = openai.chat.completions.create(
 model="nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 messages=[{"role": "user", "content": "Hello"}],
)

print(chat_completion.choices[0].message.content)
print(chat_completion.usage.prompt_tokens, chat_completion.usage.completion_tokens)

# Hello! It's nice to meet you. Is there something I can help you with, or would you like to chat?
# 11 25
```

#### Conversations

To create a longer chat-like conversation you just have to add each response message and each of the user messages to every request. This way the model will have the context and will be able to provide better answers. You can tweak it even further by providing a system message.

```python
# Assume openai>=1.0.0
from openai import OpenAI

# Create an OpenAI client with your deepinfra token and endpoint
openai = OpenAI(
 api_key="$DEEPINFRA_TOKEN",
 base_url="https://api.deepinfra.com/v1/openai",
)

chat_completion = openai.chat.completions.create(
 model="nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 messages=[
 {"role": "system", "content": "Respond like a michelin starred chef."},
 {"role": "user", "content": "Can you name at least two different techniques to cook lamb?"},
 {"role": "assistant", "content": "Bonjour! Let me tell you, my friend, cooking lamb is an art form, and I'm more than happy to share with you not two, but three of my favorite techniques to coax out the rich, unctuous flavors and tender textures of this majestic protein. First, we have the classic \"Sous Vide\" method. Next, we have the ancient art of \"Sous le Sable\". And finally, we have the more modern technique of \"Hot Smoking.\""},
 {"role": "user", "content": "Tell me more about the second method."},
 ],
)

print(chat_completion.choices[0].message.content)
print(chat_completion.usage.prompt_tokens, chat_completion.usage.completion_tokens)

# Sous le Sable! It's an ancient technique that never goes out of style, n'est-ce pas? Literally ...
# 149 324
```

The longer the conversation gets, the more time it takes the model to generate the response. The number of messages that you can have in a conversation is limited by the context size of a model. Larger models also usually take more time to respond.

 

### Streaming

Streaming any of the chat completions above is supported by adding the `stream=True` option.


```python
from openai import OpenAI

# Create an OpenAI client with your deepinfra token and endpoint
openai = OpenAI(
 api_key="$DEEPINFRA_TOKEN",
 base_url="https://api.deepinfra.com/v1/openai",
)

chat_completion = openai.chat.completions.create(
 model="nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 messages=[{"role": "user", "content": "Hello"}],
 stream=True,
)

for event in chat_completion:
 if event.choices[0].finish_reason:
 print(event.choices[0].finish_reason, event.usage["prompt_tokens"], event.usage["completion_tokens"])
 else:
 print(event.choices[0].delta.content)

# Hello
# !
# It
# 's
# nice
# ...
# 11 25
```

You can use JavaScript in the browser or in node.js to make requests

```bash
npm install openai
```

then

```javascript
import OpenAI from "openai";

const openai = new OpenAI({
 baseURL: 'https://api.deepinfra.com/v1/openai',
 apiKey: "$DEEPINFRA_TOKEN",
});

async function main() {
 const completion = await openai.chat.completions.create({
 messages: [{ role: "user", content: "Hello" }],
 model: "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 });

 console.log(completion.choices[0].message.content);
 console.log(completion.usage.prompt_tokens, completion.usage.completion_tokens);
}

main();

// Hello! It's nice to meet you. Is there something I can help you with, or would you like to chat?
// 11 25
```

#### Conversations

To create a longer chat-like conversation you just have to add each response message and each of the user messages to every request. This way the model will have the context and will be able to provide better answers. You can tweak it even further by providing a system message.

```javascript
import OpenAI from "openai";

const openai = new OpenAI({
 baseURL: 'https://api.deepinfra.com/v1/openai',
 apiKey: "$DEEPINFRA_TOKEN",
});

async function main() {
 const completion = await openai.chat.completions.create({
 messages: [
 {role: "system", content: "Respond like a michelin starred chef."},
 {role: "user", content: "Can you name at least two different techniques to cook lamb?"},
 {role: "assistant", content: "Bonjour! Let me tell you, my friend, cooking lamb is an art form, and I'm more than happy to share with you not two, but three of my favorite techniques to coax out the rich, unctuous flavors and tender textures of this majestic protein. First, we have the classic \"Sous Vide\" method. Next, we have the ancient art of \"Sous le Sable\". And finally, we have the more modern technique of \"Hot Smoking.\""},
 {role: "user", "content": "Tell me more about the second method."}
 ],
 model: "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 });

 console.log(completion.choices[0].message.content);
 console.log(completion.usage.prompt_tokens, completion.usage.completion_tokens);
}

main();

// Sous le Sable, my friend! This traditional technique hails from the ancient Mediterranean, wher...
// 149 324
```

The longer the conversation gets, the more time it takes the model to generate the response. The number of messages that you can have in a conversation is limited by the context size of a model. Larger models also usually take more time to respond.

 

### Streaming

Streaming any of the chat completions above is supported by adding the `stream: true` option.

```javascript
import OpenAI from "openai";

const openai = new OpenAI({
 baseURL: 'https://api.deepinfra.com/v1/openai',
 apiKey: "$DEEPINFRA_TOKEN",
});

async function main() {
 const completion = await openai.chat.completions.create({
 messages: [{ role: "user", content: "Hello" }],
 model: "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B",
 stream: true,
 });

 for await (const chunk of completion) {
 if (chunk.choices[0].finish_reason) {
 console.log(chunk.choices[0].finish_reason, chunk.usage.prompt_tokens, chunk.usage.completion_tokens);
 } else {
 console.log(chunk.choices[0].delta.content);
 }
 }
}

main();

// Hello
// !
// It
// 's
// nice
// ...
// 11 25
```

The [AI SDK](https://sdk.vercel.ai/) by Vercel makes it very easy to do inference.

```bash
npm install ai @ai-sdk/deepinfra
```

then

```javascript
import { createDeepInfra } from "@ai-sdk/deepinfra";
import { generateText } from "ai";

const deepinfra = createDeepInfra({
 apiKey: "$DEEPINFRA_TOKEN",
});

const { text, usage, finishReason } = await generateText({
 model: deepinfra("nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B"),
 prompt: "Write a vegetarian lasagna recipe for 4 people.",
});

console.log(text);
console.log(usage);
console.log(finishReason);
```

#### Conversations

To create a longer chat-like conversation you just have to add each response message and each of the user messages to every request. This way the model will have the context and will be able to provide better answers. You can tweak it even further by providing a system message.

```javascript
import { createDeepInfra } from "@ai-sdk/deepinfra";
import { generateText } from "ai";

const deepinfra = createDeepInfra({
 apiKey: "$DEEPINFRA_TOKEN",
});

const { text, usage, finishReason } = await generateText({
 model: deepinfra("nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B"),
 messages: [
 { role: "system", content: "Respond like a michelin starred chef." },
 {
 role: "user",
 content: "Can you name at least two different techniques to cook lamb?",
 },
 {
 role: "assistant",
 content:
 'Bonjour! Let me tell you, my friend, cooking lamb is an art form, and I\'m more than happy to share with you not two, but three of my favorite techniques to coax out the rich, unctuous flavors and tender textures of this majestic protein. First, we have the classic "Sous Vide" method. Next, we have the ancient art of "Sous le Sable". And finally, we have the more modern technique of "Hot Smoking."',
 },
 { role: "user", content: "Tell me more about the second method." },
 ],
});

console.log(text);
console.log(usage);
console.log(finishReason);
```

The longer the conversation gets, the more time it takes the model to generate the response. The number of messages that you can have in a conversation is limited by the context size of a model. Larger models also usually take more time to respond.

 

### Streaming

Streaming text responses is easy just replace `generateText` with `streamText` and read the response chunk by chunk

```javascript
import { createDeepInfra } from "@ai-sdk/deepinfra";
import { streamText } from "ai";

const deepinfra = createDeepInfra({
 apiKey: "$DEEPINFRA_TOKEN",
});

const result = streamText({
 model: deepinfra("meta-llama/Llama-3.3-70B-Instruct-Turbo"),
 prompt: "Invent a new holiday and describe its traditions.",
 system:
 "You are a professional writer. You write simple, clear, and concise content.",
});

for await (const textPart of result.textStream) {
 console.log(textPart);
}

console.log(await result.usage);
console.log(await result.finishReason);
```

It works for conversations, too.

This is an advanced and more complex API. We strongly recommend that you use OpenAI Chat Completions instead.

#### Simple prompt

To query this model you need to provide a properly formatted input string.

```bash
curl "https://api.deepinfra.com/v1/inference/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B" \
 -H "Content-Type: application/json" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -d '{
 "input": "<|im_start|>system\n<|im_end|>\n<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n<think>\n",
 "stop": [
 "<|im_end|>"
 ]
 }'
```

That will respond with

```json
{
 "request_id": "RWZDRhS5kdoM1XWwXLEshynO",
 "inference_status": {
 "status": "succeeded",
 "runtime_ms": 243,
 "cost": 0.0000436,
 "tokens_input": 12,
 "tokens_generated": 25
 },
 "results": [
 {
 "generated_text": "Hello! It's nice to meet you. Is there something I can help you with or would you like to chat for a bit?"
 }
 ],
 "num_tokens": 25,
 "num_input_tokens":12
}
```

#### Conversations

The OpenAI Chat Completions API is better suited for chat-like conversations, use it instead.

To query this model you need to provide a properly formatted input string.

However, you can still do it if you really need to. You have to add each response and each of the user prompts to every request.
You need a properly formatted input string to make it understand the current context. See the example below for some of them.
You can tweak it even further by providing a system message.

```bash
curl "https://api.deepinfra.com/v1/inference/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B" \
 -H "Content-Type: application/json" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -d '{
 "input": "<|im_start|>system\nRespond like a michelin starred chef.<|im_end|>\n<|im_start|>user\nCan you name at least two different techniques to cook lamb?<|im_end|>\n<|im_start|>assistant\n<think></think>Bonjour! Let me tell you, my friend, cooking lamb is an art form, and I'"'"'m more than happy to share with you not two, but three of my favorite techniques to coax out the rich, unctuous flavors and tender textures of this majestic protein. First, we have the classic \"Sous Vide\" method. Next, we have the ancient art of \"Sous le Sable\". And finally, we have the more modern technique of \"Hot Smoking.\"<|im_end|>\n<|im_start|>user\nTell me more about the second method.<|im_end|>\n<|im_start|>assistant\n<think>\n",
 "stop": [
 "<|im_end|>"
 ]
 }'
```

The conversation above might return something like the following

```json
{
 "request_id": "RWZDRhS5kdoM1XWwXLEshynO",
 "inference_status": {
 "status": "succeeded",
 "runtime_ms": 243,
 "cost": 0.000436,
 "tokens_input": 149,
 "tokens_generated": 338
 },
 "results": [
 {
 "generated_text": "Sous le Sable, my friend! It's an ancient technique that's been used for centuries in the Middle East and North Africa. The name itself..."
 }
 ],
 "num_tokens": 338,
 "num_input_tokens": 149
}
```

The longer the conversation gets, the more time it takes the model to generate the response. The conversation is limited by the context size of a model. Larger models also usually take more time to respond.

 

### Streaming


To do a streaming request, just pass `"stream": true`:

```bash
curl "https://api.deepinfra.com/v1/inference/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B" \
 -H "Content-Type: application/json" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -d '{
 "input": "<|im_start|>system\n<|im_end|>\n<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n<think>\n",
 "stop": [
 "<|im_end|>"
 ],
 "stream": true
 }'
```

which outputs:

```json
data: {"token": {"id": null, "text": "Hello", "logprob": 0.0, "special": false}, "generated_text": "", "details": null, "estimated_cost": null}

data: {"token": {"id": null, "text": "!", "logprob": 0.0, "special": false}, "generated_text": "", "details": null, "estimated_cost": null}

data: {"token": {"id": null, "text": " It", "logprob": 0.0, "special": false}, "generated_text": "", "details": null, "estimated_cost": null}

data: {"token": {"id": null, "text": "'s", "logprob": 0.0, "special": false}, "generated_text": "", "details": null, "estimated_cost": null}

....

data: {"token": {"id": null, "text": "", "logprob": 0.0, "special": false}, "generated_text": null, "details": {"finish_reason": "stop"}, "num_output_tokens": 25, "num_input_tokens": 12, "estimated_cost": 0.0000386}
```

#### Input format

You can see below the basic format of the input. Bear in mind that newlines often matter.

```text
<|im_start|>system
<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
<think>
```

Conversation prompts contain the history of the exchanged prompts and responses.

```text
<|im_start|>system
<|im_end|>
<|im_start|>user
First question<|im_end|>
<|im_start|>assistant
<think></think>First answer<|im_end|>
<|im_start|>user
Second question<|im_end|>
<|im_start|>assistant
<think></think>Second answer<|im_end|>
<|im_start|>user
Final question<|im_end|>
<|im_start|>assistant
<think>
```

If you want to add system prompt, it is done like this

```text
<|im_start|>system
System prompt<|im_end|>
<|im_start|>user
First question<|im_end|>
<|im_start|>assistant
<think></think>First answer<|im_end|>
<|im_start|>user
Second question<|im_end|>
<|im_start|>assistant
<think></think>Second answer<|im_end|>
<|im_start|>user
Final question<|im_end|>
<|im_start|>assistant
<think>
```

You can use our command-line tool [deepctl](/docs/getting-started) to run inferences:

This is an advanced and more complex API. We strongly recommend that you use OpenAI Chat Completions instead.

#### Simple prompt

To query this model you need to provide a properly formatted input string.

```bash
deepctl infer \
 -m 'nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B' \
 -i 'input="<|im_start|>system\n<|im_end|>\n<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n<think>\n"' \
 -i 'stop=["<|im_end|>"]'
```

That will respond with

```json
{
 "inference_status": {
 "status": "succeeded",
 "runtime_ms": 243,
 "cost": 0.0000436,
 "tokens_input": 12,
 "tokens_generated": 25
 },
 "results": [
 {
 "generated_text": "Hello! It's nice to meet you. Is there something I can help you with or would you like to chat for a bit?"
 }
 ],
 "num_tokens": 25,
 "num_input_tokens":12
}
```

#### Conversations

The OpenAI Chat Completions API is better suited for chat-like conversations, use it instead.

To query this model you need to provide a properly formatted input string.

However, you can still do it if you really need to. You have to add each response and each of the user prompts to every request.
You need a properly formatted input string to make it understand the current context. See the example below for some of them.
You can tweak it even further by providing a system message.

```bash
deepctl infer \
 -m 'nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B' \
 -i 'input="<|im_start|>system\nRespond like a michelin starred chef.<|im_end|>\n<|im_start|>user\nCan you name at least two different techniques to cook lamb?<|im_end|>\n<|im_start|>assistant\n<think></think>Bonjour! Let me tell you, my friend, cooking lamb is an art form, and I'"'"'m more than happy to share with you not two, but three of my favorite techniques to coax out the rich, unctuous flavors and tender textures of this majestic protein. First, we have the classic \"Sous Vide\" method. Next, we have the ancient art of \"Sous le Sable\". And finally, we have the more modern technique of \"Hot Smoking.\"<|im_end|>\n<|im_start|>user\nTell me more about the second method.<|im_end|>\n<|im_start|>assistant\n<think>\n"' \
 -i 'stop=["<|im_end|>"]'
```

The conversation above might return something like the following

```json
{
 "inference_status": {
 "status": "succeeded",
 "runtime_ms": 243,
 "cost": 0.000436,
 "tokens_input": 149,
 "tokens_generated": 338
 },
 "results": [
 {
 "generated_text": "Sous le Sable, my friend! It's an ancient technique that's been used for centuries in the Middle East and North Africa. The name itself..."
 }
 ],
 "num_tokens": 338,
 "num_input_tokens": 149
}
```

The longer the conversation gets, the more time it takes the model to generate the response. The conversation is limited by the context size of a model. Larger models also usually take more time to respond.

#### Input format

You can see below the basic format of the input. Bear in mind that newlines often matter.

```text
<|im_start|>system
<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
<think>
```

Conversation prompts contain the history of the exchanged prompts and responses.

```text
<|im_start|>system
<|im_end|>
<|im_start|>user
First question<|im_end|>
<|im_start|>assistant
<think></think>First answer<|im_end|>
<|im_start|>user
Second question<|im_end|>
<|im_start|>assistant
<think></think>Second answer<|im_end|>
<|im_start|>user
Final question<|im_end|>
<|im_start|>assistant
<think>
```

If you want to add system prompt, it is done like this

```text
<|im_start|>system
System prompt<|im_end|>
<|im_start|>user
First question<|im_end|>
<|im_start|>assistant
<think></think>First answer<|im_end|>
<|im_start|>user
Second question<|im_end|>
<|im_start|>assistant
<think></think>Second answer<|im_end|>
<|im_start|>user
Final question<|im_end|>
<|im_start|>assistant
<think>
```

We recommend using our NodeJS client https://github.com/deepinfra/deepinfra-node.

You can install it with

```bash
npm install deepinfra
```

#### Simple prompt

To query this model you need to provide a properly formatted input string.

```javascript
import { TextGeneration } from "deepinfra";

const DEEPINFRA_API_KEY = '$DEEPINFRA_TOKEN';
const MODEL_URL = 'https://api.deepinfra.com/v1/inference/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B';

async function main() {
 const client = new TextGeneration(MODEL_URL, DEEPINFRA_API_KEY);
 const res = await client.generate({
 "input": "<|im_start|>system\n<|im_end|>\n<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n<think>\n",
 "stop": [
 "<|im_end|>"
 ]
 });
 console.log(res.results[0].generated_text);
}

main();

// Hello! It's nice to meet you. Is there something I can help you with, or would you like to chat?
```

#### Conversations

The OpenAI Chat Completions API is better suited for chat-like conversations, use it instead.

However, you can still do it if you really need to. You have to add each response and each of the user prompts to every request.
You need a properly formatted input string to make it understand the current context. See the example below for some of them.
You can tweak it even further by providing a system message.

```javascript
import { TextGeneration } from "deepinfra";

const DEEPINFRA_API_KEY = '$DEEPINFRA_TOKEN';
const MODEL_URL = 'https://api.deepinfra.com/v1/inference/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B';

async function main() {
 const client = new TextGeneration(MODEL_URL, DEEPINFRA_API_KEY);
 const res = await client.generate({
 "input": "<|im_start|>system\nRespond like a michelin starred chef.<|im_end|>\n<|im_start|>user\nCan you name at least two different techniques to cook lamb?<|im_end|>\n<|im_start|>assistant\n<think></think>Bonjour! Let me tell you, my friend, cooking lamb is an art form, and I'm more than happy to share with you not two, but three of my favorite techniques to coax out the rich, unctuous flavors and tender textures of this majestic protein. First, we have the classic \"Sous Vide\" method. Next, we have the ancient art of \"Sous le Sable\". And finally, we have the more modern technique of \"Hot Smoking.\"<|im_end|>\n<|im_start|>user\nTell me more about the second method.<|im_end|>\n<|im_start|>assistant\n<think>\n",
 "stop": [
 "<|im_end|>"
 ]
 });
 console.log(res.results[0].generated_text);
}

main();

// Sous le Sable! It's an ancient technique that never goes out of style, n'est-ce pas? Literally ...
```

The longer the conversation gets, the more time it takes the model to generate the response.
The number of messages that you can have in a conversation is limited by the context size of a model.
Larger models also usually take more time to respond.

#### Input format

You can see below the basic format of the input. Bear in mind that newlines often matter.

```text
<|im_start|>system
<|im_end|>
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
<think>
```

Conversation prompts contain the history of the exchanged prompts and responses.

```text
<|im_start|>system
<|im_end|>
<|im_start|>user
First question<|im_end|>
<|im_start|>assistant
<think></think>First answer<|im_end|>
<|im_start|>user
Second question<|im_end|>
<|im_start|>assistant
<think></think>Second answer<|im_end|>
<|im_start|>user
Final question<|im_end|>
<|im_start|>assistant
<think>
```

If you want to add system prompt, it is done like this

```text
<|im_start|>system
System prompt<|im_end|>
<|im_start|>user
First question<|im_end|>
<|im_start|>assistant
<think></think>First answer<|im_end|>
<|im_start|>user
Second question<|im_end|>
<|im_start|>assistant
<think></think>Second answer<|im_end|>
<|im_start|>user
Final question<|im_end|>
<|im_start|>assistant
<think>
```

input

maximum length of the newly generated generated text.If explicitly set to None it will be the model's max context length minus input length or 65536, whichever is smaller

max_new_tokens

temperature to use for sampling. 0 means the output is deterministic. Values greater than 1 encourage more diversity

temperature

Sample from the set of tokens with highest probability such that sum of probabilies is higher than p. Lower values focus on the most probable tokens.Higher values sample more low-probability tokens

top_p

Float that represents the minimum probability for a token to be considered, relative to the probability of the most likely token. Must be in [0, 1]. Set to 0 to disable this.

min_p

Sample from the best k (number of) tokens. 0 means off

top_k

repetition penalty. Value of 1 means no penalty, values greater than 1 discourage repetition, smaller than 1 encourage repetition.

repetition_penalty

Up to 16 strings that will terminate generation immediately

stop

Number of output sequences to return. Incompatible with streaming

num_responses

Optional nested object with "type" set to "json_object"

response_format

Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

presence_penalty

Positive values penalize new tokens based on how many times they appear in the text so far, increasing the model's likelihood to talk about new topics.

frequency_penalty

A unique identifier representing your end-user, which can help monitor and detect abuse. Avoid sending us any identifying information. We recommend hashing user identifiers.

user

Seed for random number generator. If not provided, a random seed is used. Determinism is not guaranteed.

seed

A key to identify prompt cache for reuse across requests. If provided, the prompt will be cached and can be reused in subsequent requests with the same key.

prompt_cache_key

The webhook to call when inference is done, by default you will get the output in the response of your inference request

webhook

Whether to stream tokens, by default it will be false, currently only supported for Llama 2 text generation models, token by token updates will be sent over SSE

stream

Type

JsonObjectResponseFormat

Name

Schema

JsonSchema

JSON schema for structured output when type is 'json_schema'

JsonSchemaResponseFormat

Regex pattern for structured output when type is 'regex'

Regex

RegexResponseFormat

TextResponseFormat

Frequency Penalty

Input

Max New Tokens

Min P

Num Responses

Presence Penalty

Prompt Cache Key

Repetition Penalty

Response Format

Seed

Stop

Stream

Temperature

Top K

Top P

User

Webhook

TextGenerationIn

Generated Text

GeneratedText

estimated cost billed for the request in USD

Cost

Output Length

Runtime Ms

Status

Tokens Generated

Tokens Input

InferenceReplyStatus

Object containing the status of the inference request

Num Input Tokens

number of generated tokens, excluding prompt

Num Tokens

Request Id

Results

TextGenerationOut

The service tier used for processing the request. 'priority' processes the request with higher priority (premium rate); 'flex' processes it at lower priority for a discount, served only when spare capacity exists and may be retried/timed out under load. Both apply only to models that support the respective tier.

service_tier

model

conversation messages: (user,assistant,tool)*,user including one system message anywhere

messages

whether to stream the output via SSE or return the full response

What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic

An alternative to sampling with temperature, called nucleus sampling, where the model considers the results of the tokens with top_p probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered.

The maximum number of tokens to generate in the chat completion.

The total length of input tokens and generated tokens is limited by the model's context length. If explicitly set to None it will be the model's max context length minus input length or 65536, whichever is smaller.

max_tokens

up to 16 sequences where the API will stop generating further tokens

Up to 16 token IDs where the API will stop generating further tokens. Merged with the model's built-in stop tokens. Intended for private deployments.

stop_token_ids

A list of tools the model may call. Currently, only functions are supported as a tool.

tools

Controls which (if any) function is called by the model. none means the model will not call a function and instead generates a message. auto means the model can pick between generating a message or calling a function. required means the model must call a function. defined tool means the model must call that specific tool. none is the default when no functions are present. auto is the default if functions are present.

tool_choice

The format of the response. Currently, only json is supported.

Alternative penalty for repetition, but multiplicative instead of additive (> 1 penalize, < 1 encourage)

Whether to return log probabilities of the output tokens or not.If true, returns the log probabilities of each output token returned in the `content` of `message`.

logprobs

stream_options

Constrains effort on reasoning for reasoning models. Currently supported values are none, low, medium, high, and xhigh. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Setting to none disables reasoning entirely if the model supports.

reasoning_effort

reasoning

chat_template_kwargs

If set, the final assistant message is used as a prefix for the model to continue generating from, rather than starting a new turn. Only applicable when the last message in the conversation is an assistant message.