credits : 1

MiniMax H3 AI Video Generator

Create native-audio video with MiniMax H3 from a prompt, animate a first frame, guide a first-to-last-frame transition, or combine image, video, and audio references.

AI Video Generator Interface

How to Create a Video with MiniMax H3

Choose the H3 workflow that matches your source material, define the role of every reference, and review the current credit estimate before generating.

  1. Choose a Workflow: Start from text, upload a first frame with an optional last frame, or use Reference mode with images and optional video or audio guidance.
  2. Direct the Sequence: Describe subjects, action, shot order, camera movement, visual style, dialogue, music, sound effects, and the details that must stay unchanged.
  3. Generate and Review: Select 4–15 seconds and 768P or 2K, start the task, then inspect identity, motion, text, branding, audio synchronization, and scene continuity.

MiniMax H3 AI Video Generator FAQs

What is MiniMax H3?

MiniMax H3, also known as Hailuo 03, is a multimodal video model that can work from text, images, video, and audio references and can generate native-audio video.

Which MiniMax H3 workflows are available here?

This page offers text-to-video, first-and-last-frame image-to-video, and reference-to-video with up to nine images plus one optional video and one optional audio file in the current interface.

How long can a MiniMax H3 video be?

The current integration supports output durations from 4 through 15 seconds.

Which MiniMax H3 resolutions are available?

You can select 768P or 2K. The displayed credit estimate changes with resolution and duration.

Can MiniMax H3 generate audio?

Yes. H3 can generate native audio directed by the prompt, including dialogue, music, ambience, and sound effects, although generated audio should still be reviewed for timing and quality.

Can I use first and last frames?

Yes. In Frames mode, upload a required first frame and an optional last frame, then describe the motion and transition between them.

How are reference inputs billed?

Output is billed per generated second. If Reference mode includes a video, input-video seconds are also billed; the first five input images are included and each additional image adds credits.

How should I prompt several references?

Assign a clear role to every asset, such as identity from an image, movement from a video, and voice tone from audio, then state which subjects, designs, and scene details must remain unchanged.

Current settings and limits

MiniMax H3, also known as Hailuo 03, combines text, first-and-last-frame, and multimodal reference workflows for native-audio video generation.

Workflows
Text, first/last frames, or multimodal references
Duration
4–15 seconds
Resolution
768P or 2K
Text ratios
21:9, 16:9, 4:3, 1:1, 3:4, or 9:16
Reference input
Up to 9 images, plus one video and one audio file in this interface
Audio
Native audio can be directed through the prompt

Model result Gallery

No public result is stored for this Hub yet. Use the examples below to create the first model-specific result.

Model-specific prompt examples

Native-audio product film

Example prompt

A compact silver speaker rotates on a black glass pedestal. Move from a macro grille detail to a wide hero shot, synchronize each light pulse with a deep electronic beat, preserve the logo and product proportions, premium studio commercial.

First-to-last-frame transition

Example prompt

Transition naturally from the empty daytime storefront in the first frame to the illuminated evening launch event in the last frame. Add guests gradually, keep the architecture and signage fixed, and build the ambient crowd sound with the scene.

Reference-led character sequence

Example prompt

Use the portrait for character identity, the reference video for restrained body movement and camera rhythm, and the audio reference for vocal tone. Keep the face, jacket, and background palette consistent across the full sequence.

Know before you generate

Reference roles must be explicit

When several assets are supplied, state what each image, video, or audio file should control to reduce conflicts and drift.

Inputs affect credits

Reference-video seconds are billed with output seconds, and images beyond the first five add input-image credits. Review the displayed estimate before generating.

Review production details

Faces, hands, small text, logos, timing, audio synchronization, and unchanged areas still need human review before publishing.