DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Wan-AI/
$0.20 / second
*Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving.

Prompt
Text prompt describing the video content. Required unless media is provided. Use 'Image n' / 'Video n' identifiers to reference assets in the media array (in declaration order; images and videos counted separately).. (Default: empty)
Media
List of input media. Required unless prompt is provided. The reference_image / reference_video / reference_audio / file / link types are mutually exclusive with the first_frame / last_frame types.
Audio
Whether to generate an audio track. Default true
You need to log in to use this model
Log InSettings
Resolution
Resolution tier of the generated video (480P, 720P or 1080P). Default 1080P
Ratio
Aspect ratio of the generated video. Default 'adaptive', which picks the ratio based on the input media
Duration
Duration of the generated video in seconds (2-30). Use -1 to let the model pick the duration. Default 5. When media contains a reference_video, the input and output durations must total at most 30 seconds. (Default: empty)
Prompt Extend
Whether to enable prompt rewriting for better quality. Default true
Watermark
Whether to add AI Generated watermark. Default false
Seed
Random seed for reproducibility (Default: empty, 0 ≤ seed ≤ 2147483647)
Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.
© 2026 DeepInfra. All rights reserved.