Prompt field guide

MiniMax H3 prompt guide: structure, camera, sound, and examples

A useful MiniMax H3 prompt describes what changes over time—not only what the first frame looks like. Use the official timeline fields, name camera movement precisely, keep speakers consistent, and separate physical sound from background music. The eight templates below are original starting points you can copy, edit, and send to the Hailuo03AI generator.

Hailuo03AI currently runs text-to-video and first-frame image-to-video. MiniMax documents additional H3 modes; they are explained here for prompt literacy but are not presented as available in this product.

The three-field H3 prompt structureOfficial format
integrated_multimodal_description:
[Shot 1] style + composition + subject + action
[Shot 2] At 00:05.000, cut + camera + result

overall_soundscape:
ambience + physical sounds + non-verbal sounds

non_diegetic_music:
instruments + tempo + dynamics, or N/A

Write for time, not a still image

What a strong MiniMax H3 prompt controls

Start with the visible and audible timeline. Every detail should map to a shot, an action, a camera decision, dialogue, or a sound that can happen during the clip. Specific, compatible instructions are more useful than a long pile of visual adjectives.

Chronological shot plan

Open with [Shot 1] and add later cuts with increasing timestamps such as [Shot 2] At 00:05.000. A cut should reveal new information; use camera movement when only framing changes.

Camera type, range, and speed

Write camera motion as part of the action: push in, pull out, pan, truck, tilt, arc, track, hold static, or use POV. Add small or large amplitude and slow or fast speed only when they matter.

Sound in the right field

Put dialogue and shot-synchronized sound in the timeline, summarize ambience and physical sounds under overall_soundscape, and reserve non_diegetic_music for music the characters cannot hear. Use N/A when no score is wanted.

Five official prompt modes

Choose the H3 structure that matches your inputs

The opening instruction changes with the assets you provide. Do not paste reference labels into a plain text-to-video request, and do not describe a first-frame image as though the model has never seen it.

T2VA

Text-to-video begins directly with the three core fields and builds the complete audiovisual timeline from text.

Available here

I2VA

Image-to-video first anchors Picture 1 at 0.00 seconds, then preserves its subjects, composition, lighting, and objects while describing forward motion.

Available here

FL2VA

First-and-last-frame video aligns two pictures to the opening and ending timestamps, then describes a continuous, physically plausible path between them.

Official guide only

L2VA

Last-frame video infers a compatible opening state and gradually converges on the supplied final picture at the exact ending time.

Official guide only

Ref2VA

Full-reference video defines subjects and picture, video, and audio labels before describing what is preserved, transferred, edited, or reused across the target timeline.

Official guide only

Copy, then make it yours

MiniMax H3 prompt examples for eight common shots

These are original templates written in MiniMax's published structure, not copied showcase prompts and not claims of guaranteed output. Replace subjects, actions, timing, dialogue, and sound to fit one coherent clip before you spend credits.

T2VA10s16:9

Cinematic product reveal

A two-shot landscape product film with controlled rotation, macro texture, rain ambience, and a restrained electronic score.

integrated_multimodal_description: [Shot 1] Live-action product film, a medium-wide shot frames a matte-black wireless speaker on a wet stone plinth at blue hour. Fine rain beads on the metal grille while a narrow amber light travels across its edge. The camera arcs clockwise with small amplitude at slow speed as the speaker turns one quarter rotation, keeping the logo sharp and readable. [Shot 2] At 00:06.000, the camera cuts to a macro close-up of water droplets vibrating with the bass, then pulls out slowly as the product settles at the center of the frame.

overall_soundscape: Light rain taps against stone while a low, clean bass pulse makes the water tremble. A soft mechanical turntable hum remains underneath.

non_diegetic_music: A restrained electronic beat at a moderate tempo with deep sub-bass and sparse metallic percussion, fading cleanly at the end.
Use this prompt
T2VA10s16:9

Two-person dialogue scene

Stable speaker IDs, language-tagged dialogue, a motivated cut, and environmental sound for a short dramatic exchange.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot frames two sisters waiting beneath the awning of a closed train station at night. Rain falls beyond the warm pool of light. The older sister with a low, steady voice (S1) looks toward the empty tracks and says: <d>[English] We missed the last one.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of the younger sister with a quick, bright voice (S2). She lifts a bicycle key, smiles, and replies: <d>[English] Then we take the long way home.</d>

overall_soundscape: Rain strikes the metal awning, distant traffic passes behind the station, and the bicycle key gives a small metallic jingle.

non_diegetic_music: Sparse piano notes at a slow tempo, joined by a soft sustained cello after the second line.
Use this prompt
T2VA10s16:9

Tracking action sequence

A fast lateral tracking move, one informative cut, physically linked action, and synchronized impact sounds.

integrated_multimodal_description: [Shot 1] Live-action sports commercial, a low tracking shot follows a mountain biker accelerating along a narrow forest trail after rain. The rear tire throws small arcs of mud while the rider leans into a left turn. The camera trucks beside the bicycle at fast speed, maintaining the rider in the right third of the frame. [Shot 2] At 00:06.500, the camera cuts to a front three-quarter close shot as the rider clears a shallow stream, lands firmly, and exits toward a bright opening between the trees.

overall_soundscape: Tires grind over wet gravel, the chain clicks under load, water splashes on the landing, and the rider breathes sharply beneath the wind.

non_diegetic_music: Fast hand percussion and a short distorted bass pattern build through the jump, then stop on the landing.
Use this prompt
T2VA5s1:1

Five-second macro food shot

One simple action, a small camera move, material detail, and no background music in a square composition.

integrated_multimodal_description: [Shot 1] Live-action macro food cinematography, an extreme close-up frames a spoon breaking through the caramelized top of a small crème brûlée. The camera pushes in with small amplitude at slow speed as the brittle sugar shell cracks into irregular amber pieces and the pale custard folds around the spoon. Warm side light reveals steam and fine texture; the dessert remains centered against a dark neutral background.

overall_soundscape: The sugar crust gives a crisp crack, followed by the soft scrape of a metal spoon against ceramic and quiet room tone.

non_diegetic_music: N/A
Use this prompt
T2VA10s9:16

Vertical fashion film

A 9:16 walk-and-pause sequence with disciplined lighting, garment motion, framing, and echoing footsteps.

integrated_multimodal_description: [Shot 1] Live-action vertical fashion film, a full-body shot frames a model in a structured red coat walking through a concrete gallery. Hard morning light creates long rectangular shadows across the floor. The camera tracks backward at the model's pace while the coat hem moves naturally with each step. [Shot 2] At 00:06.000, the shot cuts to a tight profile as the model stops beside a mirrored wall, turns toward her reflection, and adjusts one cuff without looking at the camera.

overall_soundscape: Firm footsteps echo through the gallery while fabric shifts softly and distant city noise enters through an open doorway.

non_diegetic_music: A minimal drum-machine rhythm at a steady moderate tempo with one dry synth note repeating every two beats.
Use this prompt
I2VA5sadaptive

Animate a portrait first frame

Preserve identity and composition while adding only a breath, an eye-line change, and a small natural expression.

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the person shown in <Picture 1> keeps the same facial features, hairstyle, clothing, lighting direction, and background composition. The camera holds a static close shot as the person takes a quiet breath, shifts their gaze from the window toward the camera, and forms a small natural smile. Hair and loose fabric move only slightly in the existing breeze; no new objects enter the frame.

overall_soundscape: Soft room tone continues with a faint breeze and one quiet breath.

non_diegetic_music: N/A
Use this prompt
I2VA5sadaptive

Animate a product first frame

Keep labels and geometry stable, limit the rotation, and explicitly prevent unsupported new props from entering the frame.

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action product cinematography, the product in <Picture 1> keeps its exact shape, materials, label text, color, position, and background. The camera arcs clockwise with small amplitude at slow speed while the existing highlight travels naturally across the surface. The product rotates no more than fifteen degrees, then returns to a stable hero angle with the label facing the camera. Do not add hands, packaging, liquid, smoke, or extra props.

overall_soundscape: Quiet studio room tone with a subtle mechanical turntable hum.

non_diegetic_music: A single low synth note rises gently and fades before the final frame.
Use this prompt
I2VA10sadaptive

Animate a landscape first frame

Move only elements already present in the image while preserving time of day, geography, palette, and composition.

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Cinematic landscape, the mountains, lake, shoreline, clouds, and color palette in <Picture 1> remain consistent. The camera pushes forward with small amplitude at slow speed above the existing shoreline. Ripples travel across the lake, the visible grass bends gently in the wind, and the existing clouds drift gradually from left to right. Preserve the time of day and introduce no people, buildings, animals, or weather effects that are absent from the first frame.

overall_soundscape: A light breeze moves through grass while small waves touch the shore and a distant bird calls once.

non_diegetic_music: Sustained soft strings at a slow tempo, remaining quiet and even throughout.
Use this prompt

Before you generate

A five-point H3 prompt check

  • Keep the action achievable in 5, 10, or 15 seconds; shorten the story before adding more cuts.
  • Give every speaking person one stable ID such as (S1), and keep only the spoken words inside the <d> block.
  • Use exact timestamps only for later cuts; Shot 1 begins without a timestamp.
  • For image-to-video, preserve what the upload already establishes and describe the motion that develops from it.
  • Remove conflicting camera commands, duplicate sound instructions, and decorative adjectives that do not change the shot.

Sources and evidence boundary

The structure and mode definitions come from MiniMax's public H3 prompt-writing materials and API documentation. The templates on this page are editorial examples by Hailuo03AI. They have not been presented as official prompts or independently verified H3 outputs.

MiniMax H3 prompt FAQ

Questions people ask before writing an H3 prompt

What is the best MiniMax H3 prompt structure?

MiniMax's published base structure uses integrated_multimodal_description for the ordered visual and audible timeline, overall_soundscape for ambience and physical sounds, and non_diegetic_music for audience-only background music. Start each shot with composition and action, identify later cuts with increasing timestamps, and write camera motion as a natural part of the scene instead of stacking disconnected keywords.

How long should a MiniMax H3 prompt be?

There is no single ideal word count. The official API accepts prompts up to 7,000 characters, but length alone does not improve a result. Use enough detail to cover the complete timeline without contradictions. A focused five-second shot may need only one well-specified shot, while dialogue, multiple cuts, or reference relationships require more explicit timing and continuity.

How do I write a MiniMax H3 image-to-video prompt?

Begin with the official first-frame alignment instruction, then treat the uploaded image as the actual frame at 0.00 seconds. Preserve the subject's identity, clothing, objects, lighting, and spatial relationships before describing what moves next. Avoid re-inventing the whole image or adding unseen objects unless that change is deliberate and physically plausible within the short clip.

Can a MiniMax H3 prompt include dialogue and sound?

The official format includes dialogue, diegetic sound, overall soundscape, and non-diegetic music. Give each speaker a stable ID, keep the original dialogue language inside a tagged <d> block, and say whether a voice is on-screen or off-screen. Hailuo03AI has not independently published a controlled H3 audio-quality benchmark, so treat these fields as instructions rather than guaranteed output claims.

Do camera movement keywords work in H3 prompts?

MiniMax documents camera moves such as push in, pull out, pan, truck, tilt, arc, tracking, static shot, shake, POV, and roll. A useful instruction combines the motion with the subject and, when necessary, its amplitude and speed. Choose one compatible move for a moment; several simultaneous or contradictory camera commands make the intended composition less clear.