A lightweight image-to-video model that turns static visuals into controllable dynamic short clips
Seedance 1.0 Lite I2V 250428 is ByteDance's lightweight image-to-video version, suitable for creating short clips from existing product images, illustrations, or scene images. It uses an image to establish the starting point of the visual, then uses text to describe actions and camera movements, helping create dynamic assets and preview shots. Compared with the same-generation Lite T2V, it starts from existing visual designs rather than constructing visuals from text alone.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Image-driven video generation, combined with text descriptions of actions and camera movements
Available resolutions
480p、720p、1080p
Video duration
2–12 seconds
Aspect ratios
16:9、4:3、1:1、3:4、9:16、21:9、adaptive
Text input
Up to 1000 characters per text content item
Camera control
Text-based camera movement descriptions, camerafixed fixed-camera option, seed random seed
Audio and delivery
No native audio generation; obtain a video link after the task is complete
This platform calls this model via POST /seedance/videos, offering 480p, 720p, and 1080p resolutions and 2–12 second duration settings. The aspect ratios and control options listed on this page are for use on this platform and do not represent native specifications shared by all Seedance family versions; newer audio generation, multimodal reference, and video editing capabilities do not apply to this model.
Core Capabilities
Start Creating from an Existing Image
Use an image as the starting point for a video, ideal for assets with composition, color, and subject design already completed. Focus the text on what happens next, such as a gently moving hem, rising steam, or a slowly pulling-back camera, rather than redescribing the entire image, bringing static designs into a dynamic production workflow.
Express Motion and Camera Movement Separately
Prompts can specify subject motion and camera movement separately, for example, “the person gently turns their head while the camera remains fixed,” or “the subject stays still while the camera slowly pushes in.” When a fixed viewpoint is needed, use camerafixed to focus creativity on local motion rather than changing all visual elements at once.
Adapt to Different Short-Video Delivery Formats
Provides landscape, portrait, square, and adaptive aspect ratio options, and supports creating short videos at different resolutions. In the workflow, first check motion direction and composition, then adjust delivery settings. Asynchronous tasks and callbacks make it easy to integrate generation steps into asset management systems and retrieve video links by task record after completion.
Use Cases
Dynamic Product Still Assets
Input a finished product image and describe subtle motion in lighting, steam, water surfaces, or backgrounds to generate short shots suitable for advertising edits. Organize prompts around one clear action, check product appearance during review, and add brand text and key selling-point captions in post-production.
Illustration and Scene Atmosphere Shorts
Use illustrations, concept scenes, or cover images as input, and describe motion such as drifting clouds, swaying leaves, or gently moving clothing and hair to deliver video assets for openings or presentations. The original image establishes the visual design, while text adds changes over time, making it suitable for exploring different dynamic expressions of the same static work.
Storyboard Motion Previsualization
Turn finalized storyboard images into short videos, separately testing locked shots, push-ins, or pull-backs to check whether the motion serves the narrative. The deliverable is a shot previsualization for discussion rather than precise photographic simulation; save the original image, prompts, and task results so the team can compare different approaches.
How to choose this model
Choose I2V with an image; choose T2V with text only
If the subject appearance, scene, and visual style have already been determined in an image, choosing this I2V version better matches the creative starting point; if there are no visual assets yet and you want to build the scene directly from text, choose the Lite T2V version from the same date. The two variants are designed for different input methods, so you should not simply replace the model while keeping the same asset organization method.
Choose short shots and multimodal tasks separately
This model is suitable for short-shot tasks that start from an image and do not require synchronized audio. When generating videos with sound, consider Seedance 1.5 Pro; when character images, audio, or video references are needed, choose a 2.x version that supports these inputs. To edit or extend existing videos, use the corresponding 2.5 workflow.
Get started
Organize content and assets
Provide text and a role=first_frame image; put the image address in image_url.url, rather than writing it directly as a string.
Select the correct version and shot settings
Specify model=doubao-seedance-1-0-lite-i2v-250428 for /seedance/videos; first test with duration=5, resolution=720p, and a clearly defined aspect ratio. The 1.0 series does not generate native audio.
Save the final video and task records
First obtain the task_id asynchronously, then query /seedance/tasks or receive a callback; after completion, check the subject, motion, and ending, and save the selected final video and task records. This model produces no native audio, so add voiceover and music in post-production when sound is needed.
Trial suggestion: simple animation for an illustration
Input and objective
Use an illustration as the first frame: the character gently raises their head, leaves sway in the breeze, the camera remains fixed, and the artistic style and composition are preserved.
Review and next steps
Use first_frame and check the character form and subtle movements; do not port newer audio/video reference or 4K capabilities to Lite.
Usage limitations
This 1.0 Lite I2V version does not generate synchronized audio, and you cannot obtain voice-over, ambient sound, or music by enabling generate_audio. Videos requiring sound can be dubbed and mixed in post-production, or you can use a model that supports audio generation, to avoid treating silent material as a complete audiovisual deliverable.
Images here primarily serve as the starting point for the visual output and are not equivalent to the new version's character identity reference. Do not use reference_image, reference_audio, or reference_video to request cross-scene character consistency, voice reference, or motion transfer, and do not use video editing or extension parameters with this version.
Image-driven generation is still a generative process, not pixel-by-pixel animation of the original image. When product outlines, small text, or complex contact actions are involved, review the results segment by segment; it is recommended to first establish one primary action and then gradually add variations, rather than treating a single generation as a guarantee of precisely preserving every detail.
Frequently asked questions
What is the difference between this version and Lite T2V?
Lite I2V uses an image as the starting point for creation and is suitable for animating already designed visuals; Lite T2V generates visuals from text descriptions. Choose I2V first when you already have product images, illustrations, or storyboard images; choose T2V when you only have creative text. They are different task variants.
How should images and prompts be submitted?
Submit the specified model to POST /seedance/videos, and include an image_url image item and a text item in content. image_url must use an object containing url, where you can provide an image link or a Base64 data URL; the text should primarily describe the action, camera movement, and visual characteristics you want to preserve.
How long and how clear can the generated videos be?
This version supports short video durations of 2–12 seconds, with resolutions of 480p, 720p, or 1080p. It is recommended to determine the duration based on the shot objective rather than cramming all actions into one clip; for vertical or horizontal delivery, you also need to plan the aspect ratio and the subject's position in the frame.
Can it make the person in a photo speak and generate sound?
This version does not support native audio generation and cannot directly deliver talking videos with voice-over. You can first create silent moving visuals and then proceed to dubbing and editing; if you need sound to be generated at the same time, choose a version that supports audio video and separately evaluate lip-sync and content-matching results.
How do I get the video result after submission?
You can set async to true to obtain a task_id, then retrieve the completion status and video_url through task queries; you can also provide callback_url to have the completed result sent to your service. task_id is used to associate requests and results. A successful submission does not mean the video is complete; download and review it after the task is finished.