The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Maverick, a 17 billion parameter model with 128 experts

Llama-4-Maverick-17B-128E-Instruct-FP8

The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding. Llama 4 Scout, a 17 billion parameter model with 16 experts

Llama-4-Scout-17B-16E-Instruct

We introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across math, code, and reasoning tasks. 

DeepSeek-R1-Turbo

DeepSeek-R1

QwQ is the reasoning model of the Qwen series. Compared with conventional instruction-tuned models, QwQ, which is capable of thinking and reasoning, can achieve significantly enhanced performance in downstream tasks, especially hard problems. QwQ-32B is the medium-sized reasoning model, which is capable of achieving competitive performance against state-of-the-art reasoning models, e.g., DeepSeek-R1, o1-mini.

QwQ-32B

DeepSeek-V3-0324, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token, an improved iteration over DeepSeek-V3.

DeepSeek-V3-0324

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3 27B is Google's latest open source model, successor to Gemma 2

gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structured outputs and function calling. Gemma 3-12B is Google's latest open source model, successor to Gemma 2

gemma-3-12b-it

gemma-3-4b-it

Orpheus TTS is a state-of-the-art, Llama-based Speech-LLM designed for high-quality, empathetic text-to-speech generation. This model has been finetuned to deliver human-level speech synthesis, achieving exceptional clarity, expressiveness, and real-time streaming performances.

orpheus-3b-0.1-ft

Kokoro is an open-weight TTS model with 82 million parameters. Despite its lightweight architecture, it delivers comparable quality to larger models while being significantly faster and more cost-efficient. With Apache-licensed weights, Kokoro can be deployed anywhere from production environments to personal projects.

Kokoro-82M

CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs. The model architecture employs a Llama backbone and a smaller audio decoder that produces Mimi audio codes.

csm-1b

Phi-4-multimodal-instruct is a lightweight open multimodal foundation model that leverages the language, vision, and speech research and datasets used for Phi-3.5 and 4.0 models. The model processes text, image, and audio inputs, generating text outputs, and comes with 128K token context length. The model underwent an enhancement process, incorporating both supervised fine-tuning, direct preference optimization and RLHF (Reinforcement Learning from Human Feedback) to support precise instruction adherence and safety measures. The languages that each modal supports are the following: - Text: Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, Turkish, Ukrainian - Vision: English - Audio: English, Chinese, German, French, Italian, Japanese, Spanish, Portuguese

Phi-4-multimodal-instruct

DeepSeek-R1-Distill-Llama-70B is a highly efficient language model that leverages knowledge distillation to achieve state-of-the-art performance. This model distills the reasoning patterns of larger models into a smaller, more agile architecture, resulting in exceptional results on benchmarks like AIME 2024, MATH-500, and LiveCodeBench. With 70 billion parameters, DeepSeek-R1-Distill-Llama-70B offers a unique balance of accuracy and efficiency, making it an ideal choice for a wide range of natural language processing tasks. 

DeepSeek-R1-Distill-Llama-70B

DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. 

DeepSeek-V3

Llama 3.3-70B Turbo is a highly optimized version of the Llama 3.3-70B model, utilizing FP8 quantization to deliver significantly faster inference speeds with a minor trade-off in accuracy. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.

Llama-3.3-70B-Instruct-Turbo

Llama 3.3-70B is a multilingual LLM trained on a massive dataset of 15 trillion tokens, fine-tuned for instruction-following and conversational dialogue. The model is designed to be helpful, safe, and flexible, with a focus on responsible deployment and mitigating potential risks such as bias, toxicity, and misinformation. It achieves state-of-the-art performance on various benchmarks, including conversational tasks, language translation, and text generation.

Llama-3.3-70B-Instruct

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local deployment.  The model achieves 81% accuracy on the MMLU benchmark and performs competitively with larger models like Llama 3.3 70B and Qwen 32B, while operating at three times the speed on equivalent hardware.

Mistral-Small-24B-Instruct-2501

DeepSeek R1 Distill Qwen 32B is a distilled large language model based on Qwen 2.5 32B, using outputs from DeepSeek R1. It outperforms OpenAI's o1-mini across various benchmarks, achieving new state-of-the-art results for dense models.  Other benchmark results include:  AIME 2024: 72.6 | MATH-500: 94.3 | CodeForces Rating: 1691.

DeepSeek-R1-Distill-Qwen-32B

Phi-4 is a model built upon a blend of synthetic datasets, data from filtered public domain websites, and acquired academic books and Q&A datasets. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning.

phi-4

Meta developed and released the Meta Llama 3.1 family of large language models (LLMs), a collection of pretrained and instruction tuned generative text models in 8B, 70B and 405B sizes

Meta-Llama-3.1-70B-Instruct

Meta-Llama-3.1-8B-Instruct

Meta-Llama-3.1-405B-Instruct

Meta-Llama-3.1-8B-Instruct-Turbo

Meta-Llama-3.1-70B-Instruct-Turbo

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). It has significant improvements in code generation, code reasoning and code fixing. A more comprehensive foundation for real-world applications such as Code Agents. Not only enhancing coding capabilities but also maintaining its strengths in mathematics and general competencies.

Qwen2.5-Coder-32B-Instruct

Llama-3.1-Nemotron-70B-Instruct is a large language model customized by NVIDIA to improve the helpfulness of LLM generated responses to user queries. This model reaches Arena Hard of 85.0, AlpacaEval 2 LC of 57.6 and GPT-4-Turbo MT-Bench of 8.98, which are known to be predictive of LMSys Chatbot Arena Elo.  As of 16th Oct 2024, this model is #1 on all three automatic alignment benchmarks (verified tab for AlpacaEval 2 LC), edging out strong frontier models such as GPT-4o and Claude 3.5 Sonnet.

Llama-3.1-Nemotron-70B-Instruct

Qwen2.5 is a model pretrained on a large-scale dataset of up to 18 trillion tokens, offering significant improvements in knowledge, coding, mathematics, and instruction following compared to its predecessor Qwen2. The model also features enhanced capabilities in generating long texts, understanding structured data, and generating structured outputs, while supporting multilingual capabilities for over 29 languages.

Qwen2.5-72B-Instruct

The Llama 90B Vision model is a top-tier, 90-billion-parameter multimodal model designed for the most challenging visual reasoning and language tasks. It offers unparalleled accuracy in image captioning, visual question answering, and advanced image-text comprehension. Pre-trained on vast multimodal datasets and fine-tuned with human feedback, the Llama 90B Vision is engineered to handle the most demanding image-based AI tasks.  This model is perfect for industries requiring cutting-edge multimodal AI capabilities, particularly those dealing with complex, real-time visual and textual analysis.

Llama-3.2-90B-Vision-Instruct

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual question answering, bridging the gap between language generation and visual reasoning. Pre-trained on a massive dataset of image-text pairs, it performs well in complex, high-accuracy image analysis.  Its ability to integrate visual understanding with language processing makes it an ideal solution for industries requiring comprehensive visual-linguistic AI applications, such as content creation, AI-driven customer service, and research.

Llama-3.2-11B-Vision-Instruct

At 8 billion parameters, with superior quality and prompt adherence, this base model is the most powerful in the Stable Diffusion family. This model is ideal for professional use cases at 1 megapixel resolution

sd3.5

Black Forest Labs' latest state-of-the art proprietary model sporting top of the line prompt following, visual quality, details and output diversity.

FLUX-1.1-pro

FLUX.1 [schnell] is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions. This model offers cutting-edge output quality and competitive prompt following, matching the performance of closed source alternatives. Trained using latent adversarial diffusion distillation, FLUX.1 [schnell] can generate high-quality images in only 1 to 4 steps. 

FLUX-1-schnell

FLUX.1-dev is a state-of-the-art 12 billion parameter rectified flow transformer developed by Black Forest Labs. This model excels in text-to-image generation, providing highly accurate and detailed outputs. It is particularly well-regarded for its ability to follow complex prompts and generate anatomically accurate images, especially with challenging details like hands and faces.

FLUX-1-dev

Black Forest Labs' first flagship model based on Flux latent rectified flow transformers

FLUX-pro

  At 2.5 billion parameters, with improved MMDiT-X architecture and training methods, this model is designed to run “out of the box” on consumer hardware, striking a balance between quality and ease of customization. It is capable of generating images ranging between 0.25 and 2 megapixel resolution. 

sd3.5-medium

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper "Robust Speech Recognition via Large-Scale Weak Supervision" by Alec Radford  et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting. Whisper large-v3-turbo is a finetuned version of a pruned Whisper large-v3. In other words, it's the exact same model, except that the number of decoding layers have reduced from 32 to 4. As a result, the model is way faster, at the expense of a minor quality degradation.

whisper-large-v3-turbo

Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification.

whisper-large-v3

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to those leading proprietary models.

WizardLM-2-8x22B

Distil-Whisper was proposed in the paper Robust Knowledge Distillation via Large-Scale Pseudo Labelling.  This is the third and final installment of the Distil-Whisper English series. It the knowledge distilled version of OpenAI's Whisper large-v3, the latest and most performant Whisper model to date.  Compared to previous Distil-Whisper models, the distillation procedure for distil-large-v3 has been adapted to give superior long-form transcription accuracy with OpenAI's sequential long-form algorithm.

You can use cURL or any other http client to run inferences:

```bash
curl -X POST \
    -H "Authorization: bearer $DEEPINFRA_TOKEN"  \
    -F audio=@my_voice.mp3  \
    'https://api.deepinfra.com/v1/inference/distil-whisper/distil-large-v3'
```

which will give you back something similar to:

```json
{
  "text": "",
  "segments": [
    {
      "id": 0,
      "text": "Hello",
      "start": 0.0,
      "end": 1.0
    },
    {
      "id": 1,
      "text": "World",
      "start": 4.0,
      "end": 5.0
    }
  ],
  "language": "en",
  "input_length_ms": 0,
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0
  }
}

```


You can use our command-line tool [deepctl](/docs/advanced/deepctl) to run
inferences:

```bash
deepctl infer \
    -m 'distil-whisper/distil-large-v3'  \
    -i audio=@my_voice.mp3
```

which will give you back something similar to:

```json
{
  "text": "",
  "segments": [
    {
      "id": 0,
      "text": "Hello",
      "start": 0.0,
      "end": 1.0
    },
    {
      "id": 1,
      "text": "World",
      "start": 4.0,
      "end": 5.0
    }
  ],
  "language": "en",
  "input_length_ms": 0,
  "request_id": null,
  "inference_status": {
    "status": "unknown",
    "runtime_ms": 0,
    "cost": 0.0,
    "tokens_generated": 0,
    "tokens_input": 0
  }
}

```


We recommend using our NodeJS client https://github.com/deepinfra/deepinfra-node.

You can install it with

```bash
npm install deepinfra
```

and then

```javascript
import { AutomaticSpeechRecognition } from "deepinfra";
import path from "path";
import { fileURLToPath } from 'url';


const __filename = fileURLToPath(import.meta.url);
const __dirname = path.dirname(__filename);

const DEEPINFRA_API_KEY = "$DEEPINFRA_TOKEN";
const MODEL = "distil-whisper/distil-large-v3";

const main = async () => {
  const client = new AutomaticSpeechRecognition(MODEL, DEEPINFRA_API_KEY);

  const input = {
    audio: path.join(__dirname, "audio.mp3"),
  };
  const response = await client.generate(input);
  console.log(response.text);
};

main();
```


You can POST to our OpenAI Transcriptions and Translations compatible endpoint.

# Create transcription

For a given audio file and model, the endpoint will return the **transcription object** or a **verbose transcription object**.

## Request body

- **file** (Required): The audio file object to transcribe. Supported formats are `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, and `webm`.
- **model** (Required): ID of the model to use. Only `distil-whisper/distil-large-v3` for this case. For other models, refer to [models/automatic-speech-recognition](models/automatic-speech-recognition).
- **language** (Optional): The language of the input audio. Supplying the input language in ISO-639-1 format can improve accuracy and latency.
- **prompt** (Optional): An optional text prompt to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
- **response_format** (Optional): The format of the output. Options include: `json` (default), `text`, `srt`, `verbose_json`, `vtt`.
- **temperature** (Optional): Controls the sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 make it more focused and deterministic. If set to 0, the model will adjust automatically to increase temperature as needed.
- **timestamp_granularities[]** (Optional): Specifies the timestamp granularity for transcription. Requires `response_format` to be set to `verbose_json`. Options: `word` - generates timestamps for individual words, `segment` - generates timestamps for segments. Note: There is no additional latency for segment timestamps, but generating word timestamps incurs additional latency.

## Response body

The transcription object or a verbose transcription object.

### Basic request

```bash
curl "https://api.deepinfra.com/v1/openai/audio/transcriptions" \
 -H "Content-Type: multipart/form-data" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -F file="@/path/to/file/audio.mp3" \
 -F model="distil-whisper/distil-large-v3"
```
```json
{
 "text": "Imagine the wildest idea that you've ever had, and you're curious about how it might scale to something that's a 100, a 1,000 times bigger. This is a place where you can get to do that."
}
```

### Word timestamp request

```bash
curl "https://api.deepinfra.com/v1/openai/audio/transcriptions" \
 -H "Content-Type: multipart/form-data" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -F file="@/path/to/file/audio.mp3" \
 -F model="distil-whisper/distil-large-v3" \
 -F response_format="verbose_json" \
 -F "timestamp_granularities[]=word"
```

```json
{
 "task": "transcribe",
 "language": "english",
 "duration": 8.470000267028809,
 "text": "The beach was a popular spot on a hot summer day. People were swimming in the ocean, building sandcastles, and playing beach volleyball.",
 "words": [
 {
 "word": "The",
 "start": 0.0,
 "end": 0.23999999463558197
 },
 ...
 {
 "word": "volleyball",
 "start": 7.400000095367432,
 "end": 7.900000095367432
 }
 ]
}
```

### Segment timestamp request

```bash
curl "https://api.deepinfra.com/v1/openai/audio/transcriptions" \
 -H "Content-Type: multipart/form-data" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -F file="@/path/to/file/audio.mp3" \
 -F model="distil-whisper/distil-large-v3" \
 -F response_format="verbose_json" \
 -F "timestamp_granularities[]=segment"
```

```json
{
 "task": "transcribe",
 "language": "english",
 "duration": 8.470000267028809,
 "text": "The beach was a popular spot on a hot summer day. People were swimming in the ocean, building sandcastles, and playing beach volleyball.",
 "segments": [
 {
 "id": 0,
 "seek": 0,
 "start": 0.0,
 "end": 3.319999933242798,
 "text": " The beach was a popular spot on a hot summer day.",
 "tokens": [
 50364, 440, 7534, 390, 257, 3743, 4008, 322, 257, 2368, 4266, 786, 13, 50530
 ],
 "temperature": 0.0,
 "avg_logprob": -0.2860786020755768,
 "compression_ratio": 1.2363636493682861,
 "no_speech_prob": 0.00985979475080967
 },
 ...
 ]
}
```

# Create translation

For a given audio file and model, the endpoint will return the translated text to English.

## Request body

- **file** (Required): The audio file object to translate. Supported formats are `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, and `webm`.
- **model** (Required): ID of the model to use. Only `distil-whisper/distil-large-v3` for this case. For other models, refer to [models/automatic-speech-recognition](models/automatic-speech-recognition).
- **prompt** (Optional): An optional text to guide the model's style or continue a previous audio segment. The prompt should be in English.
- **response_format** (Optional): The format of the output. Options include: `json` (default), `text`, `srt`, `verbose_json`, `vtt`.
- **temperature** (Optional): The sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit.

## Response body

The translated text to English.

### Basic request

```bash
curl "https://api.deepinfra.com/v1/openai/audio/translations" \
 -H "Content-Type: multipart/form-data" \
 -H "Authorization: Bearer $DEEPINFRA_TOKEN" \
 -F file="@/path/to/file/german.m4a" \
 -F model="distil-whisper/distil-large-v3"
```

```json
{
 "text": "Hello, my name is Wolfgang and I come from Germany. Where are you heading today?"
}
```

You can use OpenAI's Python SDK to interact with our OpenAI Transcriptions and Translations compatible endpoint.

# Create transcription

For a given audio file and model, the endpoint will return the **transcription object** or a **verbose transcription object**.

## Request body

- **file** (Required): The audio file object to transcribe. Supported formats are `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, and `webm`.
- **model** (Required): ID of the model to use. Only `distil-whisper/distil-large-v3` for this case. For other models, refer to [models/automatic-speech-recognition](models/automatic-speech-recognition).
- **language** (Optional): The language of the input audio. Supplying the input language in ISO-639-1 format can improve accuracy and latency.
- **prompt** (Optional): An optional text prompt to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
- **response_format** (Optional): The format of the output. Options include: `json` (default), `text`, `srt`, `verbose_json`, `vtt`.
- **temperature** (Optional): Controls the sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 make it more focused and deterministic. If set to 0, the model will adjust automatically to increase temperature as needed.
- **timestamp_granularities[]** (Optional): Specifies the timestamp granularity for transcription. Requires `response_format` to be set to `verbose_json`. Options: `word` - generates timestamps for individual words, `segment` - generates timestamps for segments. Note: There is no additional latency for segment timestamps, but generating word timestamps incurs additional latency.

## Response body

The transcription object or a verbose transcription object.

### Example

```python
from openai import OpenAI
client = OpenAI(
 api_key="$DEEPINFRA_TOKEN",
 base_url="https://api.deepinfra.com/v1/openai",
)

audio_file = open("speech.mp3", "rb")
transcript = client.audio.transcriptions.create(
 model="whisper-1",
 file=audio_file
)
```
```json
{
 "text": "Imagine the wildest idea that you've ever had, and you're curious about how it might scale to something that's a 100, a 1,000 times bigger. This is a place where you can get to do that."
}
```

# Create translation

For a given audio file and model, the endpoint will return the translated text to English.

## Request body

- **file** (Required): The audio file object to translate. Supported formats are `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, and `webm`.
- **model** (Required): ID of the model to use. Only `distil-whisper/distil-large-v3` for this case. For other models, refer to [models/automatic-speech-recognition](models/automatic-speech-recognition).
- **prompt** (Optional): An optional text to guide the model's style or continue a previous audio segment. The prompt should be in English.
- **response_format** (Optional): The format of the output. Options include: `json` (default), `text`, `srt`, `verbose_json`, `vtt`.
- **temperature** (Optional): The sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit.

## Response body

The translated text to English.

### Basic request

```python
from openai import OpenAI
client = OpenAI(
 api_key="$DEEPINFRA_TOKEN",
 base_url="https://api.deepinfra.com/v1/openai",
)

audio_file = open("speech.mp3", "rb")
transcript = client.audio.translations.create(
 model="whisper-1",
 file=audio_file
)
```

```json
{
 "text": "Hello, my name is Wolfgang and I come from Germany. Where are you heading today?"
}
```

You can use OpenAI's JavaScript SDK to interact with our OpenAI Transcriptions and Translations compatible endpoint.

```bash
npm install openai
```

# Create transcription

For a given audio file and model, the endpoint will return the **transcription object** or a **verbose transcription object**.

## Request body

- **file** (Required): The audio file object to transcribe. Supported formats are `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, and `webm`.
- **model** (Required): ID of the model to use. Only `distil-whisper/distil-large-v3` for this case. For other models, refer to [models/automatic-speech-recognition](models/automatic-speech-recognition).
- **language** (Optional): The language of the input audio. Supplying the input language in ISO-639-1 format can improve accuracy and latency.
- **prompt** (Optional): An optional text prompt to guide the model's style or continue a previous audio segment. The prompt should match the audio language.
- **response_format** (Optional): The format of the output. Options include: `json` (default), `text`, `srt`, `verbose_json`, `vtt`.
- **temperature** (Optional): Controls the sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 make it more focused and deterministic. If set to 0, the model will adjust automatically to increase temperature as needed.
- **timestamp_granularities[]** (Optional): Specifies the timestamp granularity for transcription. Requires `response_format` to be set to `verbose_json`. Options: `word` - generates timestamps for individual words, `segment` - generates timestamps for segments. Note: There is no additional latency for segment timestamps, but generating word timestamps incurs additional latency.

## Response body

The transcription object or a verbose transcription object.

### Example

```js
import fs from "fs";
import OpenAI from "openai";

const openai = new OpenAI({
 baseURL: 'https://api.deepinfra.com/v1/openai',
 apiKey: "$DEEPINFRA_TOKEN",
});

async function main() {
 const translation = await openai.audio.translations.create({
 file: fs.createReadStream("speech.mp3"),
 model: "distil-whisper/distil-large-v3",
 });

 console.log(translation.text);
}
main();
```
```json
{
 "text": "Hello, my name is Wolfgang and I come from Germany. Where are you heading today?"
}
```

# Create translation

For a given audio file and model, the endpoint will return the translated text to English.

## Request body

- **file** (Required): The audio file object to translate. Supported formats are `flac`, `mp3`, `mp4`, `mpeg`, `mpga`, `m4a`, `ogg`, `wav`, and `webm`.
- **model** (Required): ID of the model to use. Only `distil-whisper/distil-large-v3` for this case. For other models, refer to [models/automatic-speech-recognition](models/automatic-speech-recognition).
- **prompt** (Optional): An optional text to guide the model's style or continue a previous audio segment. The prompt should be in English.
- **response_format** (Optional): The format of the output. Options include: `json` (default), `text`, `srt`, `verbose_json`, `vtt`.
- **temperature** (Optional): The sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic. If set to 0, the model will use log probability to automatically increase the temperature until certain thresholds are hit.

## Response body

The translated text to English.

### Basic request

```js
import fs from "fs";
import OpenAI from "openai";

const openai = new OpenAI({
 baseURL: 'https://api.deepinfra.com/v1/openai',
 apiKey: "$DEEPINFRA_TOKEN",
});

async function main() {
 const translation = await openai.audio.translations.create({
 file: fs.createReadStream("speech.mp3"),
 model: "distil-whisper/distil-large-v3",
 });

 console.log(translation.text);
}
main();
```

```json
{
 "text": "Hello, my name is Wolfgang and I come from Germany. Where are you heading today?"
}
```

audio

task

optional text to provide as a prompt for the first window.

initial_prompt

temperature

language that the audio is in; uses detected language if None; use two letter language code (ISO 639-1) (e.g. en, de, ja)

language

chunk_level

chunk_length_s

The webhook to call when inference is done, by default you will get the output in the response of your inference request

webhook

Audio

Chunk Length S

Chunk Level

Initial Prompt

Language

Task

Temperature

Webhook

AutomaticSpeechRecognitionIn

estimated cost billed for the request in USD

Cost

Runtime Ms

Status

Tokens Generated

Tokens Input

InferenceReplyStatus

Avg Logprob

Compression Ratio

confidence of the segment (Only in whisper-timestamped model)

Confidence

end location in input in seconds from start

No Speech Prob

Seek

start location in input in seconds from start

Start

Text

Tokens

a list of timestamped words in a segment (Only in whisper-timestamped model)

Words

Segment

Word

Object containing the status of the inference request

Inference Status

Input Length Ms

Request Id

Segments

AutomaticSpeechRecognitionOut

model

file

An optional text to guide the model's style or continue a previous audio segment.

prompt

response_format

The sampling temperature, between 0 and 1. Higher values produce more creative results.

An array specifying the granularity of timestamps to include in the transcription. Possible values are 'segment', 'word'.

distil-whisper/distil-large-v3

Input

Output