A lightweight video model for text-based creativity and short-film previsualization
Seedance 1.0 Lite T2V 250428 is ByteDance's lightweight text-to-video model. Starting from text descriptions, it transforms subjects, actions, scenes, and camera concepts into short videos. It is suitable for creative exploration without reference images, advertising shot drafts, and content previsualization, offering multiple resolution and aspect ratio options; compared with the Lite I2V version, its creative focus is on building visuals from text rather than animating existing images.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Text-to-video; fixed model doubao-seedance-1-0-lite-t2v-250428
Output resolution
480p, 720p, 1080p
Video duration
2–12 seconds, set in whole seconds
Common aspect ratios
16:9, 9:16, 1:1, 4:3, 3:4, 21:9; default 16:9
Text input
Use type=text in content; each text item supports up to 1000 characters
Creative controls
Random seed seed, fixed camera camerafixed
Result delivery
Retrieve video links through asynchronous task queries or callbacks; native audio generation is not supported
The values above are the supported invocation ranges for this model on this platform; Seedance 1.0 series capabilities do not represent the independent specifications of every variant.
Core Capabilities
Create dynamic visuals with text
No starting image is required. You can directly describe the subject's appearance, environment, actions, and lighting to form a visual draft for a short video. Prompts should focus on one main action, then add camera movement and style—for example, a person walking into a rainy nighttime street while the camera slowly follows—making it easier to compare how different concepts are presented.
Organize creation around the camera
When creating, describe the subject's actions and camera movement separately: first write how the person or object changes, then describe a push-in, follow shot, or static observation. When a fixed camera position is needed, use camerafixed to reduce conflicting camera-motion requirements in the prompt, so that each clip carries a clear visual-expression task.
Adjust visuals for delivery goals
The same idea can be produced in landscape, portrait, or square compositions, with different resolutions selected for preview and delivery. Verify the action and composition first, then increase the output setting to keep revisions focused on the content itself; when switching aspect ratios, the subject's position should also be rewritten to avoid copying a landscape composition directly into portrait format.
Use Cases
Advertising shot drafts
Enter the product category, usage environment, lighting, and display action to generate atmospheric shots or visual proposals for advertisements. This is suitable for discussing scenes and pacing when no footage is available, delivering short-video drafts for team selection; when the real product's appearance must be preserved accurately, switch to an image-driven creation method.
Short-video visual assets
Based on a content theme, describe character actions, environmental changes, or abstract imagery to create short clips suitable for portrait-oriented content. You can generate opening, transition, and ending visuals separately, then move into editing to add subtitles, voice-over, and music, rather than requesting a complete sound program in one generation.
Storyboards and concept previs
Break a script into clear shot tasks, entering the scene, subject action, and viewing perspective for each segment to generate dynamic storyboards that are easy to discuss. This is suitable for comparing different camera movements and art directions, delivering a set of shot candidates; at the previs stage, focus on whether the expression works rather than treating drafts as strictly continuous final footage.
How to choose this model
Choose T2V for text-based creation, I2V for image-based creation
When you only have copy, scene concepts, or storyboard descriptions, choose this fixed Lite T2V version and build visuals directly from text. If you already have product images, illustrations, or scene images and want to use their composition as the starting point, choose doubao-seedance-1-0-lite-i2v-250428. These are different task variants, and you should not replace model selection simply by adding an image field.
Choose a version based on audio and editing needs
Tasks suited to this model are text-driven, short-duration creations intended to deliver visual assets. When you need to generate audio simultaneously, consider Seedance 1.5 Pro or 2.x models that support audio; when you need to edit or extend existing video, choose Seedance 2.5 that supports the corresponding workflow. Do not treat these capabilities as additional switches for Lite T2V.
Getting started
Organize content and assets
Provide only type=text content items, clearly describing the shots and actions; when image-driven generation is needed, choose the corresponding Lite I2V model.
Select the correct version and shot settings
Specify model=doubao-seedance-1-0-lite-t2v-250428 for /seedance/videos; first test with duration=5, resolution=720p, and a clearly defined aspect ratio. The 1.0 series does not generate native audio.
Save the final video and task record
For asynchronous tasks, first obtain the task_id, then query /seedance/tasks or receive a callback; after completion, check the subject, actions, and ending, then save the selected final video and task record. This model outputs no native audio; add voiceover and music in post-production when audio is needed.
Trial suggestion: lightweight text storyboard
Input and goal
A paper airplane takes off from a desk, passes by a window, then lands beside a bookshelf, in a hand-drawn animation style, one smooth continuous shot, no subtitles.
Acceptance criteria and next steps
Start using only text content and plan a 5-second action; this model does not become Lite I2V by adding an image, nor does it generate native audio.
Usage Limitations
This model generates visual video and does not support automatically obtaining sound through generate_audio. Dialogue, music, and sound effects must be added in post-production; if sound and visuals must be generated together, choose a model that supports video with audio from the start of creation rather than adding audio parameters after completion.
Each video segment should be planned for 2–12 seconds and is not suitable for compressing long scripts or large numbers of actions into a single request. More complex stories should be split into shots, generated separately, and then edited together; when cross-segment character continuity is involved, appearance and scene transitions need to be checked manually, and you should not assume that all segments will be strictly consistent.
T2V builds visuals from text and is not intended to precisely reproduce product photos or lock the composition of an existing image. Actions, style, and camera requirements in prompts should remain consistent; use I2V when a real image is needed as the visual starting point, and use the corresponding editing model for video editing and extension.
Frequently Asked Questions
Are 250428 and Lite I2V the same model?
No. The full ID here refers to the fixed Lite T2V version, intended for text-to-video generation; the Lite I2V version of the same release is intended for image-to-video generation. Keep the full model ID when calling it, and choose the variant based on the source material; do not treat them as arbitrary input modes of the same model.
How should prompts be written for Lite T2V?
It is recommended to organize text by subject, scene, main action, camera, and style: first clarify what happens in the image, then describe how it should be filmed. Each text item may contain up to 1000 characters. Focus on clear and consistent action requirements, and avoid simultaneously requiring a fixed camera, orbiting shots, and rapid perspective changes.
Can it generate 4K videos or videos longer than 12 seconds?
This model offers 480p, 720p, and 1080p output options, with durations of 2–12 seconds. For longer content, generate segments and edit them together; if the task itself requires a longer single-segment output or 4K, choose another Seedance model that explicitly supports those specifications.
Can it make characters speak and automatically generate background music?
This model does not support native audio generation, and dialogue, music, or sound effects cannot be obtained by enabling generate_audio. You can first generate visual clips and then add voice-over and sound editing; if you want sound to be generated together with the video, choose 1.5 Pro or an applicable 2.x model that supports audio capabilities.
How do I submit a task and retrieve the video?
Submit the full model ID and content containing text items to POST /seedance/videos, and set duration, aspect ratio, and resolution as needed. With async, you can first obtain a task_id and then query the task result; you can also set callback_url to receive a notification upon completion and obtain data.video_url.