What is MiniMax H3?
MiniMax H3, also known as Hailuo 03, is a multimodal video model that can work from text, images, video, and audio references and can generate native-audio video.
Which MiniMax H3 workflows are available here?
This page offers text-to-video, first-and-last-frame image-to-video, and reference-to-video with up to nine images plus one optional video and one optional audio file in the current interface.
How long can a MiniMax H3 video be?
The current integration supports output durations from 4 through 15 seconds.
Which MiniMax H3 resolutions are available?
You can select 768P or 2K. The displayed credit estimate changes with resolution and duration.
Can MiniMax H3 generate audio?
Yes. H3 can generate native audio directed by the prompt, including dialogue, music, ambience, and sound effects, although generated audio should still be reviewed for timing and quality.
Can I use first and last frames?
Yes. In Frames mode, upload a required first frame and an optional last frame, then describe the motion and transition between them.
How are reference inputs billed?
Output is billed per generated second. If Reference mode includes a video, input-video seconds are also billed; the first five input images are included and each additional image adds credits.
How should I prompt several references?
Assign a clear role to every asset, such as identity from an image, movement from a video, and voice tone from audio, then state which subjects, designs, and scene details must remain unchanged.