We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Wan-AI logo

Wan-AI/

Wan3.0-Video

Partner

$0.20 / second

*

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving.

Public
Wan-AI/Wan3.0-Video cover image

Input

Prompt

Text prompt describing the video content. Required unless media is provided. Use 'Image n' / 'Video n' identifiers to reference assets in the media array (in declaration order; images and videos counted separately).. (Default: empty)

Media

List of input media. Required unless prompt is provided. The reference_image / reference_video / reference_audio / file / link types are mutually exclusive with the first_frame / last_frame types.

Audio

Whether to generate an audio track. Default true

You need to log in to use this model

Log In

Settings

Resolution

Resolution tier of the generated video (480P, 720P or 1080P). Default 1080P

Ratio

Aspect ratio of the generated video. Default 'adaptive', which picks the ratio based on the input media

Duration

Duration of the generated video in seconds (2-30). Use -1 to let the model pick the duration. Default 5. When media contains a reference_video, the input and output durations must total at most 30 seconds. (Default: empty)

Prompt Extend

Whether to enable prompt rewriting for better quality. Default true

Watermark

Whether to add AI Generated watermark. Default false

Seed

Random seed for reproducibility (Default: empty, 0 ≤ seed ≤ 2147483647)

Output

Model Information

Wan3.0-Video is an all-in-one video generation model unified support for multiple creative capabilities, including reference, editing, replication, and driving. It generates videos up to 30 seconds with omni-modal reference, and can parse files, web pages and complex images. With production-grade character consistency and lifelike visuals and sound, it delivers an immersive audiovisual experience.