All models

maestro

maestroVideo
0.6 Credits
Get your API key
maestro

From natural-language ideas to multilingual videos with subtitles

Maestro is a video production service centered on an AI director. After describing the topic, audience, and communication goals, it organizes scripts, visuals, voiceovers, music, subtitles, and rendering to turn ideas into complete videos. It is suitable for educational explainers, product promotion, and multilingual content creation, and can also incorporate image, video, and audio assets to continue editing or extending existing projects.

MaestroModel brand
VideoModel type
VideoTask capability

Specifications and API features

Clarify capacity, inputs and outputs, and calling methods before choosing a model.

Creation input
Natural-language prompt; images, videos, and audio URLs can be attached, up to 20 items
Target duration
5–300 seconds, 30 seconds by default
Aspect ratio options
9:16、16:9、1:1; 9:16 by default
Multilingual production
Up to 4 languages per request; zh-cn by default, with the first as the primary language
Video types
auto、narrated、captions、avatar、drama
Creation actions
generate、remix、edit、extend
Delivery method
Asynchronous task; retrieve the finished subtitled video and each language version upon completion

The above are the creation and API specifications for the Maestro video production entry point; the target duration does not equal the exact final video duration.

Core capabilities

Learn what maestro can bring to your work.

Organize ideas into complete videos

Maestro does more than generate visuals: it connects topics, scripts, narration, music, subtitles, and editing into a production workflow. Prompts can clearly specify who the audience is, what needs to be explained, how the opening should capture attention, and how the ending should conclude, keeping creation focused on communication goals and making it suitable for starting production directly from a content brief.

One set of visuals, multilingual expression

Specify languages through langs, with the first serving as the primary language. Other languages reuse the same visuals before corresponding voiceovers and rendering are completed. Chinese, English, Japanese, and other versions can be organized in the same task, reducing the work of repeatedly preparing visual content; delivery results are separated by language for easier individual publishing.

Continue creating from existing projects

There is no need to start from scratch after a video is completed. remix preserves the theme while adjusting the presentation, edit handles local changes such as titles, voiceovers, or color schemes, and extend expands the content. Provide a historical task ID and clear modification requirements to start a new iteration task and progressively refine the same video project.

Applicable Scenarios

Start with specific tasks to find where the model can be effective.

Knowledge Explainers and Tutorial Shorts

Enter a concept, tutorial key points, or organized article content, specify the audience's background and the conclusions you want to retain, then choose narrated to create an explainer video. The deliverable includes visuals, narration, and subtitles, making it suitable for turning text content into easy-to-watch short videos; key terms and required information should be specified in the prompt.

Product Promotion and Brand Content

Attach reference assets such as product images and logos, describe the selling points, audience, and desired visual character, then choose a landscape, portrait, or square aspect ratio. Use presets such as modern and luxury to shape the look and feel, and pair them with a narration voice to support the message; after the video is created, continue modifying the title or voice-over to produce different promotional expressions.

Existing Video Processing and International Distribution

When adding subtitles to existing material, use captions and submit the source video URL; when cross-language distribution is needed, select the target language in langs. The former focuses on processing existing videos, while the latter focuses on multilingual delivery of the same visual content, and they can be used respectively for asset organization and multi-region release of product introductions.

How to Choose This Model

Choose based on task complexity, input materials, and expected results.

Choose Maestro When You Need a Complete Video

If the task requires not only visuals but also a script, voice-over, music, and subtitles, Maestro's director-style workflow is better suited to the goal of complete delivery. Rather than making requests only around a single visual segment, prioritize describing the content structure and audience expectations here. Use generate for new projects; for projects already created, choose an action based on whether you want to reinterpret, make local edits, or continue writing.

Arrange Assets and Controls by Video Type

Choose narrated for explainer content, captions for adding subtitles to existing videos, avatar for talking-head videos, and drama for character dialogue stories. The type determines the form of expression, style adjusts the visual look and feel, and voice adjusts the narration voice; do not treat all three as a single control. If you want the system to organize the format itself, you can keep auto and clearly state the creative goal.

Get Started

From a small-scale task to formal integration.

01

Prepare the Task and Materials

Define the goal, required inputs, and output requirements, using real business examples as a starting point.

02

Try It in the API Testing Area

Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.

03

Integrate According to the API Documentation

Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.

Usage Limits

Before using it officially, understand the output quality and capability scope.

  • The target duration must be between 5–300 seconds, with a maximum of 4 languages per request and up to 20 reference URLs. Multilingual production reuses the same set of visuals; this does not mean the visual content will be redesigned for each language version. If different regions require different visuals, create separate production tasks.
  • captions requires a source video, and avatar requires a portrait; you cannot select only the type while omitting required assets. Reference content should be submitted as media URLs through file_urls; this differs from directly uploading local files. Prepare accessible asset URLs before production.
  • Maestro uses asynchronous production. A successful submission only means the task has been accepted, not that the final video is complete. Major content adjustments may trigger regeneration; visual style and narration voice are creative controls and should not be treated as tools that lock visuals frame by frame or guarantee voice-over results word for word.

Frequently Asked Questions

Answers to common questions about using maestro.

What is the difference between Maestro and tools that only generate video visuals?

Maestro is designed to organize complete video production. In addition to visuals, it includes scripts, narration, music, captions, editing, and rendering. Prompts should describe the content goal and presentation structure, rather than only the appearance of shots; it is better suited for creative tasks that move from a brief to a finished video.

How do I use my own product images, videos, or audio?

Put media URLs in file_urls and explain the purpose of each asset in the prompt, such as using product images for display, logos for brand identification, and videos for caption processing. You can submit up to 20 items; when selecting captions, provide a source video, and when selecting avatar, provide a portrait.

How many languages can be produced at once? Will the visuals differ?

You can specify up to 4 languages per request, with Chinese as the default; the first language in langs is the primary language. The remaining languages reuse the same visuals, with separate voice-overs and rendering, and language versions are provided upon completion. If you want each region to use different visual storytelling, submit separate creative requirements for each.

Should I choose remix, edit, or extend to modify a finished video?

If you want to keep the theme but change the presentation, choose remix; to change specific elements such as the title, voice-over, or color scheme, choose edit; to continue expanding existing content, choose extend. All three require ref_task_id. Describe the modification goal in the prompt, and a new task ID will then be returned.

How do I get the final video after submission?

After calling POST /maestro/videos, first obtain task_id, then use POST /maestro/tasks to check progress and results. The task goes through planning and production stages; once successful, you can obtain the finished video and each language version. Applications should distinguish between submission successful, in production, completed, and failed states.

Model information · Updated: 2026-10-01. For call parameters and billing rules, see the API and pricing sections.

Use maestro for your next task

Start with a clear goal and evaluate whether it fits your work based on real results.