Text Creation
Provide non-empty text in content, clearly specifying the subject, action, scene, and shot objective, and use one short video to express one clear task.
MiniMax-H3-Max uses unified content input and supports integer durations of 5–15 seconds. First select input materials and an output tier based on the task, then retrieve the completed result through platform task queries.
Provide non-empty text in content, clearly specifying the subject, action, scene, and shot objective, and use one short video to express one clear task.
Use first frames, last frames, or multimodal reference materials as documented. The type and role of different materials determine their purpose; do not use legacy top-level fields such as prompt and image_urls.
Schedule tasks in whole seconds within the 5–15 second range, choose a duration that matches the number of subject actions, and review the output after completion.
Use text or keyframes to discuss camera movement, composition, and pacing, providing dynamic reference for subsequent production.
Use existing images as supported reference inputs, describe presentation actions and scene atmosphere, and generate short assets for later editing.
When combining image, video, or audio materials, first clarify the reference purpose of each asset, then submit according to content combination limits.
H3 Max provides 480P and 768P. For 2K, choose a model that supports that tier rather than submitting 2K to H3 Max.
Reference images and reference videos affect the cost structure. Refer to the Pricing page for current billing rules and package pricing.
Explicitly specify model=MiniMax-H3-Max, and provide text and any required reference materials in content.
Choose 480P or 768P and an integer duration of 5–15 seconds. ratio requirements depend on the input workflow; configure it according to the API documentation.
Use async=true to save task_id, then query via /minimax/tasks, or use callback_url to receive completion notifications.
Submit a request to POST /minimax/videos, explicitly set model=MiniMax-H3-Max, and provide non-empty text in content.
480P, 768P, and integer durations of 5–15 seconds are supported; 2K is not supported for this model.
Use the content items and role values defined in the documentation to indicate first-frame, last-frame, or reference use, and comply with quantity and combination limits for each material type.
Audio input incurs no additional charge, and the first two images are free; additional images and reference videos are billed according to their respective input usage, while output video costs are calculated under current rules.
An asynchronous request first returns task_id. Query the same task, and after confirming task.status=succeeded, retrieve task.content.url; if it fails, read the task error information.
Updated: 2026-10-01. Specifications and fees are subject to the current API and pricing pages.
This is only a basic call example. See the full docs for more parameters and advanced usage.
View Docs