We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

FastVideo/

FastWan2.2-TI2V-5B-FullAttn-Diffusers

$0.0225 per 5s video (720p); 480p half price

A fast, high-quality text-to-video model — 5-second clips at 24 fps in 720p or 480p, landscape or portrait, from a text prompt. A 3-step DMD distillation of Wan2.2-TI2V-5B by FastVideo (Hao AI Lab).

FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers cover image

Input

Prompt

text prompt describing the video content. This model rewards dense, specific prompts: name the subject and its motion, the camera move, the lighting, and the style.

Seconds

Clip duration: always 5 seconds (fixed/required for this model).

Resolution

Output resolution class. 720p renders 1280x704 (704 on the short side: the Wan2.2 VAE needs multiples of 32, the industry-standard convention for this model family); 480p renders 832x480.

Orientation

Output orientation: landscape (1280x704) or portrait (704x1280).

You need to log in to use this model

Log In

Settings

Seed

specify a seed for reproducible output (Default: empty)

Output

Model Information

FastWan2.2-TI2V-5B

From the FastVideo team (Hao AI Lab):

We're excited to introduce the FastWan2.2 series — a new line of models finetuned with our novel Sparse-distill strategy. This approach jointly integrates DMD and VSA in a single training process, combining the benefits of both distillation to shorten diffusion steps and sparse attention to reduce attention computations, enabling even faster video generation.

FastWan2.2-TI2V-5B-FullAttn-Diffusers is built upon Wan-AI/Wan2.2-TI2V-5B-Diffusers. It supports efficient 3-step inference and produces high-quality videos at 121×704×1280 resolution. For training, we used simulated forward for the generator model, making the process data-free. The current model is trained using only DMD.

Links from the authors: Project page · GitHub · VSA paper

On DeepInfra

  • Resolutions: 720p-class (1280×704) and 480p (832×480), each in landscape or portrait
  • Clips: 5 seconds, 121 frames at 24 fps
  • Speed: a clip typically renders in ~1.5–3.5 seconds
  • Text-to-video only

Prompting tips

This model rewards dense, specific prompts. The strongest results name, in order: the camera move, the subject and its action, the environment, the lighting, and the style — see the example prompt on the model page. Water, weather, light effects, animals and nature scenes render especially well.

Things to know:

  • negative_prompt has no effect: the model is distilled to run without classifier-free guidance, so there is no negative pass to steer. Put what you want in the prompt instead.
  • Motion-speed instructions ("slow motion", "fast") are followed loosely — a side effect of few-step distillation.
  • Legible text (signs, labels) is not a strength; fine turbulent detail (smoke, fire) can move unnaturally fast.
  • The number of denoising steps is fixed at 3 by the distillation — it is not a quality dial.

Licensing

Model weights are released by FastVideo under Apache-2.0.