DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

bosonai logo

bosonai/

HiggsAudioV2.5

$20.00

/ 1M characters

HiggsAudioV2.5 is a high-quality neural text-to-speech (TTS) model designed for natural-sounding voice generation across a wide range of use cases. It focuses on clarity, stable prosody, and consistent pacing, making it suitable for both short prompts and longer narration.

Public
bosonai/HiggsAudioV2.5 cover image
api

Input

Input text

Text to convert to speech

You need to log in to use this model

Log In

Settings

ServiceTier

The service tier used for processing the request. 'priority' processes the request with higher priority (premium rate); 'flex' processes it at lower priority for a discount, served only when spare capacity exists and may be retried/timed out under load. Both apply only to models that support the respective tier. For compatibility, 'auto' is treated as 'priority' and 'standard_only' as 'default'.

Fail Fast

If true, the request is rejected immediately with HTTP 429 when the model has no spare capacity, instead of waiting in the queue. Opt-in; the default (false) keeps standard queueing behavior.

Voice

Voice name to use. Can be a preset voice (belinda, chadwick, etc.) or a custom voice ID for voice cloning.

Response Format

Output format (only pcm supported). (Default: pcm)

Stream

Whether to stream audio bytes in chunks

Output

Waiting for audio data... Submit request to start streaming.