A director-style video model that brings static images to life according to shot intent
minimax-i2v-director is the image-to-video director mode in the MiniMax Hailuo series, designed for creators who already have reference images and want greater control over action and camera expression. It uses a first-frame image and prompts to organize dynamic storytelling, focusing on reducing motion randomness, improving instruction following, and delivering cinematic shots. It is suitable for transforming product images, character visuals, and storyboard sketches into video assets.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Image-to-video director mode, starting creation from a reference image
Input
First-frame image URL (first_image_url) and text prompt (prompt)
Native control focus
Precise motion control, prompt adherence, and preset camera settings
API endpoint
POST /hailuo/videos;model=minimax-i2v-director;action=generate
Task mode
Supports async asynchronous submission and callback_url result callbacks
Result delivery
Video link video_url, with task ID and status information
Native director capabilities are reflected in motion and camera expression; this platform uses first-frame images, prompts, and task result management.
Core capabilities
Extend motion from the first-frame composition
Use a reference image to establish the subject, scene, and starting composition, then describe the desired action in text. The creative focus thus shifts from “what the image should look like” to “how the image should move,” making it suitable for projects with existing visual assets and convenient for testing different shot approaches around the same first frame.
Make motion serve the narrative
The focus of Director is reducing motion randomness so that actions align more closely with creative intent. Prompts can separately specify subject behavior and camera movement, such as a character slowly turning their head while the camera pushes forward; expressing intentions separately makes it easier to organize clear dynamic storytelling than stacking abstract style terms.
Highlight cinematic camera expression
The native design of I2V-01-Director combines preset camera settings with enhanced prompt adherence, emphasizing the role of shots in storytelling. When using it, describe the image around changes in shot scale, visual focus, and movement direction, treating a clear shot intention as the generation goal rather than merely asking for a lively image.
Applicable Scenarios
Turn Product Images into Showcase Shots
Provide a product image with the appearance and composition already confirmed, then describe camera intentions such as a slow push-in and keeping the subject stable to generate video assets for presentation. Suitable for exploring product presentation approaches before filming; after delivery, packaging text, labels, and details should still be checked to ensure they meet brand usage requirements.
Character Storyboard Motion Previsualization
Use a character design image or storyboard frame as the first frame, clearly specify character actions, gaze direction, and camera movement, and generate motion previsualizations for discussion. Directors, designers, and editors can use these to discuss shot pacing before deciding which actions to retain and which compositions need to be redesigned.
Visual Asset Generation Workflow
The application can submit image links together with camera descriptions, use asynchronous tasks to track generation progress, and retrieve video links for previewing or editing after completion. Suitable for connecting a design asset library to the video creation workflow, so that every generation can be associated with the source image, prompt, and corresponding task result.
How to Choose This Model
Already Have an Image and Care About How It Moves
minimax-i2v and minimax-i2v-director both start from a reference image; Director is explicitly positioned to enhance creative control. If the main goal is to add motion to static assets, start with a standard image-to-video workflow; if camera direction, action planning, and prompt adherence are key review criteria, Director mode is worth prioritizing rather than judging all visual quality metrics by version name.
Distinguish Image-to-Video and Text-to-Video by Creative Starting Point
I2V-01-Director and T2V-01-Director both belong to the Director model series, but their creative starting points differ. When the visual identity has already been determined, image-to-video mode makes it easier to organize subsequent actions around the first frame; when there is only a text concept, choose a text-to-video model. minimax-i2v-director should not be treated as a general text-to-video entry point that requires no image, and specifications from other Hailuo versions should not be applied to it.
Get Started
Prepare the Creative Starting Point for This Model
Prepare a publicly accessible image for first_image_url, and separately describe the subject action, camera action, and elements that need to be preserved in the prompt.
Use the Hailuo Video Endpoint
Specify model=minimax-i2v-director, action=generate, and a prompt for /hailuo/videos; for image-to-video, also provide the first-frame URL, without applying the MiniMax H3 content structure.
Retrieve the Video Link
Set async=true to first obtain the task_id, query it through /hailuo/tasks or receive results via callback_url; read the successful video_url and proceed to editing only after a complete review.
Trial suggestion: clear subject and camera movement
Input and objective
Use a portrait photo as the first frame, first have the person raise their head and look out the window, then have the camera slowly pan to the right, maintaining the scene lighting and clothing.
Acceptance criteria and next steps
Describe the character's movement and the camera movement separately, and check whether both occur at the same time; Director Mode is not frame-by-frame timeline control.
Usage boundaries
Director Mode enhances creative control; it is not a frame-by-frame animation editor, nor does it guarantee that every action will be reproduced precisely. It is recommended to structure each prompt around one main action and one type of camera movement, avoiding simultaneous requests for conflicting viewpoints, movement directions, or narrative changes.
The first frame is the starting point of creation, not a promise to lock the appearance throughout. For tasks with strict requirements for character identity, product structure, text labels, and so on, inspect changes in the generated video, paying particular attention to details after movement, occlusion, and viewpoint changes, before deciding whether to use it for final delivery.
This model is suitable for generating dynamic assets from reference images and should not be used directly as a complete video-production tool. Voice-over, subtitles, editing, and audio-video synchronization should be included in subsequent production plans; do not automatically apply the high resolution, duration, or audio capabilities of other video versions to this model.
Frequently Asked Questions
What images need to be prepared for minimax-i2v-director?
Prepare a reference image to use as the starting frame, and submit the link via first_image_url. It is recommended to first determine the subject, composition, and scene, then use the prompt to describe the intended motion; this is an image-to-video director mode and should not be used as a text-to-video model without an image.
How does it differ from regular minimax-i2v?
Both create videos using reference images, but Director places greater emphasis on motion control, camera expression, and prompt adherence. If you need to compare different camera movement approaches or have specific requirements for action arrangement, try director mode first; this does not mean all image quality metrics will necessarily be higher.
How should prompts be written for director mode?
Write the subject's actions and the camera's movements separately, then add any necessary scene changes. For example, first describe a person turning toward the window, then describe the camera slowly pushing in. Keep the goal of a single shot clear whenever possible, and avoid cramming multiple scene changes, complex actions, and conflicting camera movements into the same instruction.
How do I submit a task and get the generated video?
Send a generation request to /hailuo/videos, specifying model as minimax-i2v-director and action as generate, and include the image link and prompt. You can use the async method or set a result callback, use task_id to associate the task, and then obtain video_url from the completed result.
Does Director mean that the camera can be controlled frame by frame?
No. It focuses on making generated actions and camera expression more closely follow the prompt, rather than providing frame-by-frame timelines or keyframe editing. When you need to precisely arrange every frame, complete multi-shot stitching, or synchronize voice-over, use the generated result as source material before moving into editing and post-production.