Define complete video shots with text and multimedia assets
wan3.0-video is a video generation model in the Alibaba Wan 3 series, suitable for creating shots from text concepts, first and last frame designs, or multimedia references. It brings different creative approaches together in unified asset inputs, generating videos with specified or smart durations while allowing selection of resolution, aspect ratio, and sound. Through this platform, shot drafts, product showcases, and reference-driven creations can be integrated into asynchronous workflows.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Generation methods
Text generation, first and last frame generation, multimedia reference generation
Platform duration
2–30 second integers; duration=-1 for smart duration
Platform resolution
480P、720P、1080P
Output aspect ratio
adaptive、16:9、4:3、1:1、3:4、9:16
Asset inputs
media supports seven asset types, with a maximum total of 10 items
Asynchronous task; returns video link, dimensions, and thumbnail information
The figures and controls above are the wan3.0-video invocation specifications for this platform; the Prime accelerated version is a different model.
Core capabilities
From descriptions to asset-driven creation
In addition to describing scenes with prompts, you can submit reference images, videos, and audio to translate appearance, actions, and pacing requirements into specific assets. Wan 3 identifies the creation mode based on media, making it suitable for projects with an existing visual direction, while text can focus on explaining which elements should be retained and which content should be regenerated.
Define shot boundaries with first and last frames
When the opening and closing images are already determined, you can provide first_frame and last_frame separately, then describe the motion and camera changes in between. This approach is suitable for product reveals, subject movement, and transition drafts, keeping creation centered on clearly defined start and end images; it is a different mode from reference asset generation and requests must be organized separately.
Control video according to delivery goals
Duration, resolution, aspect ratio, and sound can be selected in combination based on publishing goals. First validate shots at 480P, then create delivery versions at 720P or 1080P; landscape, portrait, and square content can each use their own aspect ratio. When sound is needed, explicitly enable audio; reference audio should be specified separately as an asset.
Use Cases
Product Showcases and Advertising Shots
Provide product images and camera movement references, and specify the subject, environment, presentation sequence, and visual characteristics to preserve in the prompt to generate product reveal or demonstration clips. If the opening and closing compositions are already designed, use first-and-last-frame mode instead to deliver video assets ready for editing and review.
Social Content and Concept Previsualization
Start with a scene description, clearly define character actions, environmental atmosphere, and camera position, then choose an aspect ratio and duration suited to the publishing channel. Create low-resolution drafts during the concept stage, then adjust output specifications after confirming the direction. This is suitable for turning copywriting ideas into watchable shot plans rather than directly producing a complete long-form video.
Reference-Driven Creative Pipeline
Submit image, video, or audio assets together with creative instructions for style exploration, action references, and pacing previsualization. The application saves the asynchronous task_id, queries results through /wan/tasks, then archives successful videos and thumbnails to the asset library, connecting manual review, editing, and publishing workflows.
How to Choose This Model
Trade-offs with Wan 2.6 Calling Methods
When a new project needs to switch between text, first-and-last-frame, and mixed references, wan3.0-video's unified media input makes assets easier to organize. Existing Wan 2.6 integrations can continue using their original task models and parameters such as action. During migration, redesign the asset structure rather than simply replacing the model name, and do not treat old parameters as the creative capabilities of the new model.
How to Choose Between Standard and Prime
wan3.0-video can be used for asset-driven creation and background generation tasks; Wan 3.0 Prime is a separate accelerated version, better suited to projects where generation wait time is the primary consideration. When choosing, test actual task completion times and final output quality. Do not interpret Prime's speed positioning as a fixed duration for the standard version, and do not judge image quality solely by the version name.
Getting Started
Assign a Role to Each Asset
In media, submit assets using types such as first_frame, last_frame, or reference_image/video/audio, clearly stating which ones constrain the subject and which provide motion or rhythm. The total array can contain up to 10 items.
Configure a Wan 3 Video Request
Explicitly select model=wan3.0-video for /wan/videos, and set duration, ratio, and resolution according to the task; audio is disabled by default, so set audio=true when sound output is needed.
Deliver Asynchronously and Validate
After saving task_id, retrieve the final video through /wan/tasks or a callback; check the actual duration, audio and video, and asset references, then save frames for the next shot.
Suggested trial: multi-asset product shot
Inputs and goal
The image defines the backpack's appearance, the video provides the walking motion, and the audio provides the rhythm; a person carries the backpack through a park, with the camera smoothly following and the motion coordinated with the rhythm.
Acceptance and next steps
Use media to distinguish reference_image, reference_video, and reference_audio; for output with sound, explicitly set audio=true, and check whether the roles of the assets are correctly expressed.
Usage boundaries
First and last frames cannot be mixed with reference assets in the same request. Choose frame mode when you need to lock the starting and ending images, and choose reference mode when you need to draw on appearance, motion, or sound; the total number of media items is limited to 10, and assets should be organized around the same creative goal to avoid stacking incompatible control methods together.
The reference video here is generation material and does not mean precise editing or seamless extension of the original footage. If the goal is to keep every existing shot unchanged and modify only local content, a dedicated editing workflow should be used; this model is better suited to recreating video clips based on prompts and references.
Successful asynchronous submission does not mean the video has already been completed. The application needs to save the task_id, check the task's final status, and retrieve the video link after success. file and link are asset input types, and you cannot assume that any file format or web page can be used directly based on this; validate the actual assets before integration.
Frequently Asked Questions
Does wan3.0-video require changing the model ID based on the creation method?
No. When calling /wan/videos, use model=wan3.0-video. Text-based creation is expressed through prompt, while first/last-frame or reference-based creation is expressed through media, and the system identifies the mode accordingly. First/last frames and reference materials should be submitted separately; do not mix them in the same request just to cover more control conditions.
What is the difference between duration=-1 and specifying a number of seconds?
You can specify an integer duration of 2–30 seconds, which is suitable for tasks with an existing editing rhythm or a clearly defined delivery length. duration=-1 lets the model intelligently choose the duration, making it suitable for initially exploring shot expression. If the video must fit a fixed placement or timeline, prioritize specifying the number of seconds to make subsequent editing easier to arrange.
Are reference audio and generated sound the same thing?
No. reference_audio is used to submit audio reference material, while audio controls whether the generated video includes sound, with a default of false. If you want the final video to include sound, explicitly enable audio and describe the sound goal in the prompt; submitting audio reference alone does not mean sound output has been enabled.
How do I retrieve a video generated by an asynchronous task?
After setting async=true, first save the returned task_id, then query the final task status through /wan/tasks. Successful videos are stored on this platform's CDN, and the result provides a video link and may include size and thumbnail information. The workflow should distinguish between submission, processing, and successful delivery, rather than only checking whether the request succeeded.
How does a reference video affect billing?
Reference videos are billed based on the combined actual seconds of input video and successfully generated output video; text, image, and audio inputs do not add video seconds. Therefore, for the same output length, usage charges may differ between creations with a reference video and text-only creations. View current pricing in Pricing; final usage is based on actual usage.