Chronological shot plan
Open with [Shot 1] and add later cuts with increasing timestamps such as [Shot 2] At 00:05.000. A cut should reveal new information; use camera movement when only framing changes.
Prompt field guide
A useful MiniMax H3 prompt describes what changes over time—not only what the first frame looks like. Use the official timeline fields, name camera movement precisely, keep speakers consistent, and separate physical sound from background music. The eight templates below are original starting points you can copy, edit, and send to the Hailuo03AI generator.
Hailuo03AI currently runs text-to-video and first-frame image-to-video. MiniMax documents additional H3 modes; they are explained here for prompt literacy but are not presented as available in this product.
integrated_multimodal_description:
[Shot 1] style + composition + subject + action
[Shot 2] At 00:05.000, cut + camera + result
overall_soundscape:
ambience + physical sounds + non-verbal sounds
non_diegetic_music:
instruments + tempo + dynamics, or N/AWrite for time, not a still image
Start with the visible and audible timeline. Every detail should map to a shot, an action, a camera decision, dialogue, or a sound that can happen during the clip. Specific, compatible instructions are more useful than a long pile of visual adjectives.
Open with [Shot 1] and add later cuts with increasing timestamps such as [Shot 2] At 00:05.000. A cut should reveal new information; use camera movement when only framing changes.
Write camera motion as part of the action: push in, pull out, pan, truck, tilt, arc, track, hold static, or use POV. Add small or large amplitude and slow or fast speed only when they matter.
Put dialogue and shot-synchronized sound in the timeline, summarize ambience and physical sounds under overall_soundscape, and reserve non_diegetic_music for music the characters cannot hear. Use N/A when no score is wanted.
Five official prompt modes
The opening instruction changes with the assets you provide. Do not paste reference labels into a plain text-to-video request, and do not describe a first-frame image as though the model has never seen it.
T2VA
Text-to-video begins directly with the three core fields and builds the complete audiovisual timeline from text.
Available hereI2VA
Image-to-video first anchors Picture 1 at 0.00 seconds, then preserves its subjects, composition, lighting, and objects while describing forward motion.
Available hereFL2VA
First-and-last-frame video aligns two pictures to the opening and ending timestamps, then describes a continuous, physically plausible path between them.
Official guide onlyL2VA
Last-frame video infers a compatible opening state and gradually converges on the supplied final picture at the exact ending time.
Official guide onlyRef2VA
Full-reference video defines subjects and picture, video, and audio labels before describing what is preserved, transferred, edited, or reused across the target timeline.
Official guide onlyCopy, then make it yours
These are original templates written in MiniMax's published structure, not copied showcase prompts and not claims of guaranteed output. Replace subjects, actions, timing, dialogue, and sound to fit one coherent clip before you spend credits.
A two-shot landscape product film with controlled rotation, macro texture, rain ambience, and a restrained electronic score.
integrated_multimodal_description: [Shot 1] Live-action product film, a medium-wide shot frames a matte-black wireless speaker on a wet stone plinth at blue hour. Fine rain beads on the metal grille while a narrow amber light travels across its edge. The camera arcs clockwise with small amplitude at slow speed as the speaker turns one quarter rotation, keeping the logo sharp and readable. [Shot 2] At 00:06.000, the camera cuts to a macro close-up of water droplets vibrating with the bass, then pulls out slowly as the product settles at the center of the frame.
overall_soundscape: Light rain taps against stone while a low, clean bass pulse makes the water tremble. A soft mechanical turntable hum remains underneath.
non_diegetic_music: A restrained electronic beat at a moderate tempo with deep sub-bass and sparse metallic percussion, fading cleanly at the end.Stable speaker IDs, language-tagged dialogue, a motivated cut, and environmental sound for a short dramatic exchange.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot frames two sisters waiting beneath the awning of a closed train station at night. Rain falls beyond the warm pool of light. The older sister with a low, steady voice (S1) looks toward the empty tracks and says: <d>[English] We missed the last one.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of the younger sister with a quick, bright voice (S2). She lifts a bicycle key, smiles, and replies: <d>[English] Then we take the long way home.</d>
overall_soundscape: Rain strikes the metal awning, distant traffic passes behind the station, and the bicycle key gives a small metallic jingle.
non_diegetic_music: Sparse piano notes at a slow tempo, joined by a soft sustained cello after the second line.A fast lateral tracking move, one informative cut, physically linked action, and synchronized impact sounds.
integrated_multimodal_description: [Shot 1] Live-action sports commercial, a low tracking shot follows a mountain biker accelerating along a narrow forest trail after rain. The rear tire throws small arcs of mud while the rider leans into a left turn. The camera trucks beside the bicycle at fast speed, maintaining the rider in the right third of the frame. [Shot 2] At 00:06.500, the camera cuts to a front three-quarter close shot as the rider clears a shallow stream, lands firmly, and exits toward a bright opening between the trees.
overall_soundscape: Tires grind over wet gravel, the chain clicks under load, water splashes on the landing, and the rider breathes sharply beneath the wind.
non_diegetic_music: Fast hand percussion and a short distorted bass pattern build through the jump, then stop on the landing.One simple action, a small camera move, material detail, and no background music in a square composition.
integrated_multimodal_description: [Shot 1] Live-action macro food cinematography, an extreme close-up frames a spoon breaking through the caramelized top of a small crème brûlée. The camera pushes in with small amplitude at slow speed as the brittle sugar shell cracks into irregular amber pieces and the pale custard folds around the spoon. Warm side light reveals steam and fine texture; the dessert remains centered against a dark neutral background.
overall_soundscape: The sugar crust gives a crisp crack, followed by the soft scrape of a metal spoon against ceramic and quiet room tone.
non_diegetic_music: N/AA 9:16 walk-and-pause sequence with disciplined lighting, garment motion, framing, and echoing footsteps.
integrated_multimodal_description: [Shot 1] Live-action vertical fashion film, a full-body shot frames a model in a structured red coat walking through a concrete gallery. Hard morning light creates long rectangular shadows across the floor. The camera tracks backward at the model's pace while the coat hem moves naturally with each step. [Shot 2] At 00:06.000, the shot cuts to a tight profile as the model stops beside a mirrored wall, turns toward her reflection, and adjusts one cuff without looking at the camera.
overall_soundscape: Firm footsteps echo through the gallery while fabric shifts softly and distant city noise enters through an open doorway.
non_diegetic_music: A minimal drum-machine rhythm at a steady moderate tempo with one dry synth note repeating every two beats.Preserve identity and composition while adding only a breath, an eye-line change, and a small natural expression.
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, the person shown in <Picture 1> keeps the same facial features, hairstyle, clothing, lighting direction, and background composition. The camera holds a static close shot as the person takes a quiet breath, shifts their gaze from the window toward the camera, and forms a small natural smile. Hair and loose fabric move only slightly in the existing breeze; no new objects enter the frame.
overall_soundscape: Soft room tone continues with a faint breeze and one quiet breath.
non_diegetic_music: N/AKeep labels and geometry stable, limit the rotation, and explicitly prevent unsupported new props from entering the frame.
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action product cinematography, the product in <Picture 1> keeps its exact shape, materials, label text, color, position, and background. The camera arcs clockwise with small amplitude at slow speed while the existing highlight travels naturally across the surface. The product rotates no more than fifteen degrees, then returns to a stable hero angle with the label facing the camera. Do not add hands, packaging, liquid, smoke, or extra props.
overall_soundscape: Quiet studio room tone with a subtle mechanical turntable hum.
non_diegetic_music: A single low synth note rises gently and fades before the final frame.Move only elements already present in the image while preserving time of day, geography, palette, and composition.
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Cinematic landscape, the mountains, lake, shoreline, clouds, and color palette in <Picture 1> remain consistent. The camera pushes forward with small amplitude at slow speed above the existing shoreline. Ripples travel across the lake, the visible grass bends gently in the wind, and the existing clouds drift gradually from left to right. Preserve the time of day and introduce no people, buildings, animals, or weather effects that are absent from the first frame.
overall_soundscape: A light breeze moves through grass while small waves touch the shore and a distant bird calls once.
non_diegetic_music: Sustained soft strings at a slow tempo, remaining quiet and even throughout.Before you generate
The structure and mode definitions come from MiniMax's public H3 prompt-writing materials and API documentation. The templates on this page are editorial examples by Hailuo03AI. They have not been presented as official prompts or independently verified H3 outputs.
MiniMax H3 prompt FAQ
MiniMax's published base structure uses integrated_multimodal_description for the ordered visual and audible timeline, overall_soundscape for ambience and physical sounds, and non_diegetic_music for audience-only background music. Start each shot with composition and action, identify later cuts with increasing timestamps, and write camera motion as a natural part of the scene instead of stacking disconnected keywords.
There is no single ideal word count. The official API accepts prompts up to 7,000 characters, but length alone does not improve a result. Use enough detail to cover the complete timeline without contradictions. A focused five-second shot may need only one well-specified shot, while dialogue, multiple cuts, or reference relationships require more explicit timing and continuity.
Begin with the official first-frame alignment instruction, then treat the uploaded image as the actual frame at 0.00 seconds. Preserve the subject's identity, clothing, objects, lighting, and spatial relationships before describing what moves next. Avoid re-inventing the whole image or adding unseen objects unless that change is deliberate and physically plausible within the short clip.
The official format includes dialogue, diegetic sound, overall soundscape, and non-diegetic music. Give each speaker a stable ID, keep the original dialogue language inside a tagged <d> block, and say whether a voice is on-screen or off-screen. Hailuo03AI has not independently published a controlled H3 audio-quality benchmark, so treat these fields as instructions rather than guaranteed output claims.
MiniMax documents camera moves such as push in, pull out, pan, truck, tilt, arc, tracking, static shot, shake, POV, and roll. A useful instruction combines the motion with the subject and, when necessary, its amplitude and speed. Choose one compatible move for a moment; several simultaneous or contradictory camera commands make the intended composition less clear.