按时间推进的镜头表
用 [Shot 1] 开场,后续切镜使用递增时间戳,例如 [Shot 2] At 00:05.000。切镜应带来新信息;只是景别变化时优先写运镜。
提示词实用指南
好用的 MiniMax H3 提示词不只描述第一帧长什么样,还要说明画面随时间如何变化。按官方时间线字段写动作,准确标明运镜,固定说话人编号,并把环境声与背景音乐分开。下面 8 个原创模板可直接复制、修改,再带入 Hailuo03AI 生成器。
Hailuo03AI 当前接入文生视频与首帧图生视频。MiniMax 官方还记录了其他 H3 模式;本页会解释其提示词结构,但不会把它们包装成本站已提供的功能。
integrated_multimodal_description:
[Shot 1] style + composition + subject + action
[Shot 2] At 00:05.000, cut + camera + result
overall_soundscape:
ambience + physical sounds + non-verbal sounds
non_diegetic_music:
instruments + tempo + dynamics, or N/A写时间变化,不是静态画面
从可见、可听的时间线开始。每个细节都应对应一个镜头、动作、运镜决定、对白或片段中真实会发生的声音。具体且不冲突的指令,比堆很多视觉形容词更有效。
用 [Shot 1] 开场,后续切镜使用递增时间戳,例如 [Shot 2] At 00:05.000。切镜应带来新信息;只是景别变化时优先写运镜。
把 Push In、Pull Out、Pan、Truck、Tilt、Arc、Tracking、Static、POV 等运镜自然写进动作。只有确实重要时,再补充幅度大小和速度快慢。
对白和镜头同步声音写进时间线;环境声、物理动作声汇总到 overall_soundscape;角色听不到、只有观众能听到的配乐写进 non_diegetic_music。不需要配乐时写 N/A。
五种官方提示词模式
不同素材对应不同开场指令。纯文生视频不要硬塞参考素材标签;首帧图生视频也不要把模型已经看到的图片当成完全未知画面重新发明。
T2VA
文生视频直接从三个核心字段开始,完全用文字搭建一条完整的音视频时间线。
本站可用I2VA
图生视频先把 Picture 1 锚定在 0.00 秒,再保留人物、构图、光线和物体关系,描述画面如何向前发展。
本站可用FL2VA
首尾帧模式把两张图分别对齐开头与结尾时间,再描述一条连续、符合物理逻辑的中间路径。
仅官方指南L2VA
末帧模式从给定结尾反推合理的开场状态,并让动作、物体与构图逐步收敛到最后一帧。
仅官方指南Ref2VA
全参考模式先定义人物与图片、视频、音频标签,再说明目标时间线里哪些内容被保留、迁移、编辑或复用。
仅官方指南先复制,再改成你的镜头
这些是按 MiniMax 公布结构编写的原创模板,不是搬运的展示 prompt,也不代表生成效果保证。花积分前,请把主体、动作、时间、对白和声音改成一条统一、可执行的短片。
横屏两镜头产品片,包含受控旋转、微距材质、雨声环境与克制的电子配乐。
integrated_multimodal_description: [Shot 1] Live-action product film, a medium-wide shot frames a matte-black wireless speaker on a wet stone plinth at blue hour. Fine rain beads on the metal grille while a narrow amber light travels across its edge. The camera arcs clockwise with small amplitude at slow speed as the speaker turns one quarter rotation, keeping the logo sharp and readable. [Shot 2] At 00:06.000, the camera cuts to a macro close-up of water droplets vibrating with the bass, then pulls out slowly as the product settles at the center of the frame.
overall_soundscape: Light rain taps against stone while a low, clean bass pulse makes the water tremble. A soft mechanical turntable hum remains underneath.
non_diegetic_music: A restrained electronic beat at a moderate tempo with deep sub-bass and sparse metallic percussion, fading cleanly at the end.用固定说话人编号、语言标签、动机明确的切镜和环境声完成一段短戏。
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot frames two sisters waiting beneath the awning of a closed train station at night. Rain falls beyond the warm pool of light. The older sister with a low, steady voice (S1) looks toward the empty tracks and says: <d>[English] We missed the last one.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of the younger sister with a quick, bright voice (S2). She lifts a bicycle key, smiles, and replies: <d>[English] Then we take the long way home.</d>
overall_soundscape: Rain strikes the metal awning, distant traffic passes behind the station, and the bicycle key gives a small metallic jingle.
non_diegetic_music: Sparse piano notes at a slow tempo, joined by a soft sustained cello after the second line.高速横向跟拍、一次有信息量的切镜、连续动作与同步落地声音。
integrated_multimodal_description: [Shot 1] Live-action sports commercial, a low tracking shot follows a mountain biker accelerating along a narrow forest trail after rain. The rear tire throws small arcs of mud while the rider leans into a left turn. The camera trucks beside the bicycle at fast speed, maintaining the rider in the right third of the frame. [Shot 2] At 00:06.500, the camera cuts to a front three-quarter close shot as the rider clears a shallow stream, lands firmly, and exits toward a bright opening between the trees.
overall_soundscape: Tires grind over wet gravel, the chain clicks under load, water splashes on the landing, and the rider breathes sharply beneath the wind.
non_diegetic_music: Fast hand percussion and a short distorted bass pattern build through the jump, then stop on the landing.一个清晰动作、小幅推进、真实材质声,并在方形构图里明确不要背景音乐。
integrated_multimodal_description: [Shot 1] Live-action macro food cinematography, an extreme close-up frames a spoon breaking through the caramelized top of a small crème brûlée. The camera pushes in with small amplitude at slow speed as the brittle sugar shell cracks into irregular amber pieces and the pale custard folds around the spoon. Warm side light reveals steam and fine texture; the dessert remains centered against a dark neutral background.
overall_soundscape: The sugar crust gives a crisp crack, followed by the soft scrape of a metal spoon against ceramic and quiet room tone.
non_diegetic_music: N/A9:16 行走与停顿片段,控制光影、衣料运动、构图以及空间脚步回声。
integrated_multimodal_description: [Shot 1] Live-action vertical fashion film, a full-body shot frames a model in a structured red coat walking through a concrete gallery. Hard morning light creates long rectangular shadows across the floor. The camera tracks backward at the model's pace while the coat hem moves naturally with each step. [Shot 2] At 00:06.000, the shot cuts to a tight profile as the model stops beside a mirrored wall, turns toward her reflection, and adjusts one cuff without looking at the camera.
overall_soundscape: Firm footsteps echo through the gallery while fabric shifts softly and distant city noise enters through an open doorway.
non_diegetic_music: A minimal drum-machine rhythm at a steady moderate tempo with one dry synth note repeating every two beats.保留身份与构图,只增加呼吸、视线变化和一个自然的小表情。
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, the person shown in <Picture 1> keeps the same facial features, hairstyle, clothing, lighting direction, and background composition. The camera holds a static close shot as the person takes a quiet breath, shifts their gaze from the window toward the camera, and forms a small natural smile. Hair and loose fabric move only slightly in the existing breeze; no new objects enter the frame.
overall_soundscape: Soft room tone continues with a faint breeze and one quiet breath.
non_diegetic_music: N/A稳定标签与几何形状,限制旋转幅度,并明确阻止画面凭空增加道具。
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action product cinematography, the product in <Picture 1> keeps its exact shape, materials, label text, color, position, and background. The camera arcs clockwise with small amplitude at slow speed while the existing highlight travels naturally across the surface. The product rotates no more than fifteen degrees, then returns to a stable hero angle with the label facing the camera. Do not add hands, packaging, liquid, smoke, or extra props.
overall_soundscape: Quiet studio room tone with a subtle mechanical turntable hum.
non_diegetic_music: A single low synth note rises gently and fades before the final frame.只移动图片中已有元素,同时保留时段、地理关系、色调与原始构图。
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Cinematic landscape, the mountains, lake, shoreline, clouds, and color palette in <Picture 1> remain consistent. The camera pushes forward with small amplitude at slow speed above the existing shoreline. Ripples travel across the lake, the visible grass bends gently in the wind, and the existing clouds drift gradually from left to right. Preserve the time of day and introduce no people, buildings, animals, or weather effects that are absent from the first frame.
overall_soundscape: A light breeze moves through grass while small waves touch the shore and a distant bird calls once.
non_diegetic_music: Sustained soft strings at a slow tempo, remaining quiet and even throughout.生成前检查
本页结构和模式定义来自 MiniMax 公开的 H3 提示词材料与 API 文档。页面中的 8 个模板由 Hailuo03AI 编辑整理,不冒充官方 prompt,也没有包装成已独立实测的 H3 输出。
MiniMax H3 提示词 FAQ
MiniMax 公布的基础结构使用 integrated_multimodal_description 编排按顺序发生的画面与声音,用 overall_soundscape 汇总环境声和物理动作声,再用 non_diegetic_music 描述只有观众能听到的背景音乐。每个镜头先交代构图与动作,后续切镜使用递增时间戳,运镜则自然写进当前场景。
没有一个固定的最佳字数。官方 API 的提示词上限是 7,000 字符,但变长本身不会提升结果。内容应刚好覆盖完整时间线,并且不互相冲突。聚焦的 5 秒单镜头可以很短;对白、多次切镜或参考素材关系,则需要更明确的时间与连续性描述。
先使用官方首帧对齐指令,把上传图片当作 0.00 秒的真实画面。描述后续动作前,先保留人物身份、服装、物体、光线和空间关系。不要把整张图重新发明一遍,也不要随意加入原图没有的物体,除非这种变化是明确、必要并且能在短片时长内合理发生。
官方格式包含对白、画内声音、整体声景与画外配乐。每个说话人使用固定编号,在带语言标签的 <d> 区块中保留原始对白,并说明声音来自画内还是画外。Hailuo03AI 尚未发布受控的 H3 音频质量实测,因此这些字段应理解为生成指令,而不是对输出效果的保证。
MiniMax 文档列出了 Push In、Pull Out、Pan、Truck、Tilt、Arc、Tracking、Static Shot、Shake、POV 和 Roll 等运镜。有效写法会把运镜和主体动作放在一起,必要时再注明幅度与速度。一个时刻只选兼容的运镜,避免同时堆叠多条互相矛盾的镜头命令。