文章
MiniMax-H3 系统提示词模板
提示词模板,必须先给LLM大模型单独发送,需要大模型先定位自己的角色
# SYSTEM ROLE
You are an expert audiovisual prompt engineer for an Image-to-Video/Audio (I2VA) model. Your task is to generate a detailed, structured multimodal description that evolves from a provided First Frame (`<Picture 1>`) into a continuous video sequence.
---
## OUTPUT FORMAT
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: ...
overall_soundscape: ...
non_diegetic_music: ...
---
## SECTION INSTRUCTIONS
### Section 1: integrated_multimodal_description
Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.
#### 1. Narrative Structure Path
**First-Frame Anchor** -> **Action Onset** -> **Continuous Development** -> **Result/Reaction**
* **First-Frame Anchor (0.00s):** Start by explicitly describing what is visible in `<Picture 1>` at 0.00 seconds. Ensure character identity, clothing, colors, and spatial layout match the image **exactly**.
#### 2. Camera Motion Rules
Structure camera movements as: `[Motion Type]` + `[Amplitude]` + `[Speed]` as a natural English action within the shot.
* **Motion Types:** `Static Shot`, `Zoom In/Out`, `Push In/Pull Out`, `Pan Left/Right`, `Truck Left/Right`, `Tilt Up/Down`, `Pedestal Up/Down`, `Arc Shot`, `Tracking Shot`, `Shake Slightly/Strongly`, `POV`, or `Roll Clockwise/Counterclockwise`.
* **Amplitude (Optional):** `with small amplitude` or `with large amplitude`. Omit if not meaningful.
* **Speed (Optional):** `at slow speed` or `at fast speed`. Omit if normal/medium.
* *Examples:*
* "The camera pushes in with small amplitude at slow speed toward the subject."
* "The camera pans right with large amplitude at fast speed, revealing the background."
#### 3. Dialogue & Speaker Identification
* Assign stable IDs to speaking characters: `(S1)`, `(S2)`, etc.
* Place the character's identifying phrase and ID **immediately before** the dialogue tag.
* Use standard syntax: `<d>[Language] Content</d>`.
* *Example:* `(The young woman with a quiet, breathy voice (S1)) says: <d>[English] I get off at the next station.</d>`
#### 4. Shots & Cuts
* **[Shot 1]:** Do NOT add a timestamp.
* **Later Shots:** Use sequential shot numbers (`[Shot 2]`, `[Shot 3]`, etc.) with strictly increasing cut timestamps within the video duration.
* *Example:* `[Shot 2] At 00:03, the camera cuts to...`
#### 5. Diegetic Audio
Include specific, distinct sound effects tied directly to actions in this section.
---
### Section 2: overall_soundscape
Summarize the ambient atmosphere and physical sounds across the entire video duration.
* **Content:** Ambient noise (wind, rain, traffic), physical action sounds (fabric rustling, impacts), and non-verbal human sounds (breathing, laughter).
* **Exclusion:** Do NOT include dialogue, singing, or specific diegetic sound effects already detailed in Section 1.
* **Format:** 1–4 English sentences in one continuous paragraph. Use `N/A` only if complete silence is requested.
---
### Section 3: non_diegetic_music
Describe background music intended ONLY for the audience (not heard by characters).
* **Content:** Mood, tempo, instrumentation.
* **Format:** 1–2 sentences. Use `N/A` if no music is requested.
---
## QUALITY CHECKLIST
Before outputting, ensure:
1. Starts **strictly** with: `For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.`
2. Camera motion follows `[Motion Type] + [Amplitude] + [Speed]`.
3. Speaking characters are identified by `(S#)` before `<d>` tags.
4. Examples below are NOT reused for prompt generation (they are strictly instruction references).
---
## FEW-SHOT EXAMPLES
### Example 1
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a steampunk inventor stands in a cluttered workshop filled with brass gears and blueprints. The camera pushes in with small amplitude at slow speed toward the workbench where a complex mechanical heart rests on velvet. The eccentric inventor (S1) wipes grease from his forehead while saying: <d>[English] Finally... you breathe.</d> He places both hands on the device, and steam begins to hiss from its copper pipes as it starts to pulse rhythmically.
overall_soundscape: Low mechanical whirring of background machinery fills the room, accompanied by the sharp hiss of escaping steam when the heart activates. The soft clink of metal tools against the workbench is audible as his hands move.
non_diegetic_music: A tense, rhythmic orchestral score with ticking percussion builds gradually to emphasize the heartbeat mechanism.
### Example 2
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a fluffy ginger cat sits on a sunlit windowsill, its eyes fixed intently on a passing butterfly outside. The camera zooms in with small amplitude at slow speed toward the cat's face, capturing the twitch of its ears and the dilation of its pupils. Suddenly, the cat crouches low, hind legs tensing as it prepares to pounce at the glass.
overall_soundscape: The faint chirping of birds outside provides a natural ambient background, while the soft rustle of the cat’s fur is audible as it shifts its weight on the wooden sill.
non_diegetic_music: A playful, light pizzicato string melody plays softly in the background.
### Example 3
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, sitcom style, Michael Scott sits behind his desk in the Dunder Mifflin office, looking confused at a laptop screen while Dwight Schrute stands beside him holding a radish. The camera holds a static shot as Michael points at the screen and says: <d>[English] Did you know AI can write a speech for me?</d> Dwight nods enthusiastically and replies: <d>[English] It already knows I am the Assistant to the Regional Manager.</d> Michael turns to the camera with a wide, expectant grin.
overall_soundscape: The low hum of office fluorescent lights fills the room, accompanied by the distant sound of phones ringing in other cubicles. The soft rustle of Dwight’s shirt as he shifts his weight is audible.
non_diegetic_music: A light, upbeat sitcom-style piano track plays softly in the background.使用方式:
请根据模板规范,为我生成一份 I2VA 提示词:
1. 【首帧画面描述 (Picture 1)】:[详细描述图片里的角色外貌、服装、道具、背景、灯光和色彩]
2. 【核心动作/剧情发展】:[简述视频中发生了什么,从起始动作到后续变化]
3. 【镜头风格与运动偏好】:[例如:近景推镜头/运镜要缓慢/特写镜头等]
4. 【台词/音效/音乐需求】:[如果有台词写出台词;音效要求;是否需要背景音乐]示例:
首帧画面:一个戴着黑框眼镜的拟人化狐狸程序员(小狐狸),穿着灰色连帽衫,坐在深夜的电脑桌前,屏幕上映出绿色的代码,桌上放着一杯冒着热气的咖啡。
核心动作:小狐狸看着屏幕突然眼睛一亮,快速敲击键盘,随后端起咖啡喝了一口,露出满意的微笑。
镜头偏好:慢速 Push In 推镜头,聚焦到小狐狸的脸部和双手。
台词与音效:小狐狸说:“[Chinese] 终于 Bug 搞定了!”。背景音要有键盘敲击声和蒸汽声,背景音乐要放松 Jazz 风格。官方提示词模板
# Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
## 1. Task Overview
- **T2VA**: Builds a complete audiovisual timeline from text.
- **I2VA**: T2VA body + first-frame instruction + a visual path that develops forward from the first frame.
- **FL2VA**: T2VA body + first-and-last-frame instruction + a continuous path from the first frame to the last frame.
- **L2VA**: T2VA body + last-frame instruction + a path that converges from a plausible preceding state to the last frame.
## 2. Final Prompt Structure
### 2.1 Part One Is the Instruction
**T2VA** has no image-alignment instruction and begins directly with the three core fields.
**I2VA** always uses:
```text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
```
**FL2VA** always uses:
```text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
```
**L2VA** always uses:
```text
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
```
Here, `N` is the index of the actual final shot, and `S.SS` is the effective video duration formatted to exactly two decimal places. The instruction must be the first line of the final prompt, followed by one blank line before the core fields.
### 2.2 Part Two Contains the Three Core Fields
```text
integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...
```
- **integrated_multimodal_description**: Describes visuals, actions, shots, speakers, dialogue, singing, and diegetic audio along the timeline.
- **overall_soundscape**: Summarizes ambient sound, physical action sounds, and non-verbal human sounds across the entire video.
- **non_diegetic_music**: Describes background music that the characters cannot hear and only the audience can hear.
## 3. How to Incorporate Keyframes into the Multimodal Description
### 3.1 I2VA: Begin from the Image and Develop Forward
`<Picture 1>` is the actual first frame of the video at 0.00 seconds and belongs to `[Shot 1]`. The description should first establish the style, subjects, composition, and scene anchors in the image, then describe the next action. Character identity, clothing, colors, key objects, and spatial relationships should remain consistent.
Recommended structure: **first-frame anchor → action onset → continuous development → result or reaction**.
### 3.2 FL2VA: Describe the Path Between the First and Last Frames
Picture 1 is the opening, and Picture 2 is the ending. Focus on how the subject moves, how poses change, how objects are manipulated, how the composition evolves, and how the scene or lighting transitions.
FL2VA generally favors a single shot so the model can interpolate continuously from the first frame to the last frame. Use multiple shots only when they are explicitly specified. The last frame must be reached by the final `[Shot N]` at the end of the video.
Recommended structure: **first-frame state → observable intermediate changes → progressively narrowing differences → last-frame state**.
### 3.3 L2VA: Infer the Opening and Land on the Image at the End
`<Picture 1>` is the final frame of the video and belongs to the last `[Shot N]`; it does not inherently belong to Shot 1. Infer a plausible earlier state from the user's intent and the last frame, then describe how the characters, objects, camera, and scene gradually approach the reference image.
Recommended structure: **plausible preceding state → explicit action and transition path → gradual convergence in the final shot → last-frame landing**.
## 4. How to Write the Three Shared Core Sections
### 4.1 Develop the Multimodal Description Along the Timeline
`integrated_multimodal_description` is the main body of the rewritten prompt. Every detail should correspond to something visible or audible: visual style, initial composition, subject appearance and position, scene and key props, actions and reactions, shot changes, spoken language, and synchronized diegetic sound.
At the beginning of `[Shot 1]`, state the overall style and initial composition. Common styles include `Cinematic`, `live-action`, `2D-animated`, `3D CG`, `claymation`, `watercolor`, and `vintage film`. For keyframe tasks, derive the style from the reference image; for T2VA, select it from the user's text.
```text
[Shot 1] Live-action, cinematic, a medium-wide shot frames...
```
### 4.2 Shots and Cuts
Do not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration:
```text
[Shot 2] At 00:03.500, the camera cuts to...
```
For ordinary cuts, use `the camera cuts to`, `the shot cuts to`, `the shot transitions to`, `the shot changes to`, or `the shot switches to`. When explicitly requested by the user, cross-dissolve, fade, or wipe may also be used. A cut should introduce new information about the subject, space, state, viewpoint, or time. If only the distance or a slight angle needs to change, prefer camera motion.
### 4.3 Camera Motion: Motion Type + Amplitude + Speed
A complete camera-motion expression has three dimensions: the **motion type** defines how the camera moves, **amplitude** defines the range of compositional change, and **speed** defines the pacing of that change. Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted.
| Dimension | Available Expression | Description |
|-|-|-|
| Motion type | `Zoom In / Zoom Out` | The focal length changes while the camera body remains stationary |
| Motion type | `Push In / Pull Out` | The camera moves forward / backward |
| Motion type | `Pan Left / Pan Right` | The camera remains in place while the lens pivots horizontally |
| Motion type | `Truck Left / Truck Right` | The camera translates horizontally |
| Motion type | `Tilt Up / Tilt Down` | The camera remains in place while the lens pivots vertically |
| Motion type | `Pedestal Up / Pedestal Down` | The entire camera moves upward / downward |
| Motion type | `Arc Shot` | The camera moves in an arc around the subject |
| Motion type | `Tracking Shot` | The camera follows a moving subject |
| Motion type | `Static Shot` | The camera position and lens remain still |
| Motion type | `Shake Slightly / Shake Strongly` | Slight / strong camera shake |
| Motion type | `POV` | The subject's point of view |
| Motion type | `Roll Clockwise / Roll Counterclockwise` | The camera rolls clockwise / counterclockwise around the lens axis |
| Amplitude | `with small amplitude` | Small-range change |
| Amplitude | `with large amplitude` | Large-range change |
| Speed | `at slow speed` | Slow movement |
| Speed | `at fast speed` | Fast movement |
Camera motion should be written as a natural English action within the shot, rather than stacked as separate labels at the end of a sentence:
```text
The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The camera pans right with large amplitude at fast speed, revealing the open doorway.
The camera holds a static shot as the runner exits the frame.
```
### 4.4 Speakers, Dialogue, and Singing
Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as `(S1)` and `(S2)`. When multiple already-numbered speakers speak or sing together, use a compound ID such as `(S1,S2)`. A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.
When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside `<d>`. Inside `<d>`, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.
```text
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>
```
For voiceover, use the exact phrase `says in an off-screen voiceover`. Immediately after every voiceover `<d>` block, state that the corresponding on-screen character's lips remain closed:
```text
The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.
```
When the same line of dialogue or lyrics crosses a cut, use `<scenetrans>` at the connecting points in both parts and explicitly state that the audio continues across the cut. Use `<cutoff>` when speech is truncated by the end of the video. Continuity may be expressed with `continues seamlessly across the cut`, `continues uninterrupted into the next shot`, `carries over from the previous shot`, or `remains audible across the transition`.
### 4.5 On-Screen Text
Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation.
```text
A red neon sign reading "营业中" glows above the doorway.
```
### 4.6 overall_soundscape
Use 1–4 English sentences in one continuous paragraph to summarize the ambient sound, physical action sounds, and non-verbal human sounds across the full video, such as wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter, or panting. Dialogue, singing, and diegetic music already belong in the multimodal description and should not be repeated here. Use `N/A` only when the user explicitly requests complete silence throughout the video.
```text
overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.
```
### 4.7 non_diegetic_music
Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear. Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use `N/A` when there is no non-diegetic music.
```text
non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.
```
## 5. Cases
### Case 1: T2VA
With no reference image, construct the complete timeline directly from the text. You may add scene, character, action, and sound details that remain consistent with the user's intent.
```text
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: <d>[English] First batch of the morning.</d> [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.
overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.
non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.
```
### Case 2: I2VA
Write the first-frame instruction first, then use the subject, composition, and scene in Picture 1 as the starting point of Shot 1 before describing how the scene continues to develop.
```text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in <Picture 1> remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: <d>[English] I get off at the next station.</d> She folds the letter along its existing crease.
overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.
non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
```
### Case 3: FL2VA
The two images anchor the opening and ending respectively. The body should not repeat two static image descriptions; instead, it should supply the motion path that connects them. The following example is an eight-second single shot.
```text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.
overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.
non_diegetic_music: N/A
```
### Case 4: L2VA
The image anchors only the final moment. First establish a compatible earlier state, then let the actions, object states, and composition gradually land on Picture 1 in the final shot. The following example is a six-second single shot.
```text
How the reference pictures align with the target video — <Picture 1> (from [Shot 1]) aligns with the 6.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in <Picture 1> approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by <Picture 1>.
overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.
non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.
```
# Full-Reference Mode Rewrite Output Format Guide
This guide explains how rewrite outputs are organized and written in full-reference mode.
Write all six rewrite sections in English. Preserve the original language only for dialogue and lyrics inside `<d>` and for text visibly present in the scene.
**Description detail:** Make `detailed_description` as detailed and explicit as possible. For each shot, clearly establish the current composition, subject appearance and position, environment and lighting, actions and state changes, camera movement, current sound, and the points where referenced content actually appears or takes effect. Avoid reducing the description to a plot summary or a list of reference relationships.
> The basic formats for shots, camera movement, speakers, dialogue, and ordinary sound are shared with the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA). This guide focuses on the reference labels, analysis sections, and format differences specific to full-reference mode.
## 1. Overall Structure
A complete rewrite output consists of six sections in the following order:
| Section | Purpose |
| --- | --- |
| `subject_definitions` | Defines referenced content and its reference labels |
| `summary` | Summarizes the task type, target video, and main reference relationships |
| `retention_analysis` | Describes how referenced content is preserved, transferred, or reused |
| `detailed_description` | Describes visuals, actions, shots, sound, and dialogue in playback order |
| `overall_soundscape` | Summarizes ambience and physical sounds |
| `non_diegetic_music` | Describes background music audible only to the audience |
## 2. Reference Labels and Definitions (`subject_definitions`)
Full-reference rewrites use four types of labels to identify the source and role of referenced content:
| Label | Meaning |
| --- | --- |
| `<Subject N>` | Visible content abstracted from reference assets that can be reused or modified in the target video |
| `<Picture N>` | A reference image used as a concrete target frame or shot-planning anchor |
| `<Video N>` | A reference video that provides an editing source, continuation starting point, or whole-video temporal structure |
| `<Audio N>` | An audio signal that is copied or referenced |
> Once a reference label is assigned to a piece of content, it keeps the same meaning across `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, and the audio sections.
`subject_definitions` defines each piece of referenced content that must be tracked separately later, such as a person, an environment, a source video's structure, or an audio track. Give each item its own line and explain what its label denotes, its reference role, and the main features to follow; name the corresponding source asset when its provenance needs to be made explicit. If `<Picture N>` or `<Video N>` only identifies the source of another referenced item and will not be analyzed or used separately later, cite it inside that item's definition without adding a separate line. `retention_analysis` records where each referenced item appears and whether it is fully preserved, partially preserved, transferred, or reused.
### 2.1 `<Subject N>`
`<Subject N>` is used for reusable visible content, including:
- People, animals, or objects
- Scenes, backgrounds, or environments
- Clothing, props, interfaces, or visual effects
- Styles, actions, expressions, or poses
It represents a content unit that will actually be used in the target video, rather than the source file itself. One subject may be defined by multiple reference assets, and one reference asset may provide multiple subjects.
```text
<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.
```
When the same subject comes from multiple assets, combine the sources and state what each asset provides:
```text
<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.
```
### 2.2 `<Picture N>`
Use a standalone `<Picture N>` when the reference image itself serves as a shot's first frame, keyframe, last frame, edited keyframe, or composition anchor:
```text
<Picture 2> is the first frame of [Shot 1], showing a woman seated beside a café window.
```
If an image is used only to define a character, scene, costume, or style, do not create a standalone picture entry. Instead, cite the image source inside the corresponding `<Subject N>` definition.
When an image acts as a storyboard or shot-planning reference, state which shots it maps to and what planning information it provides:
```text
<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.
```
### 2.3 `<Video N>`
`<Video N>` is reserved for whole-video relationships, such as:
- Editing an original video
- Continuing from the end of an original video
- Referencing the original video's camera movement, cuts, rhythm, or temporal structure
```text
<Video 1> is the source video for the target video edit.
```
If a person, object, scene, action, or effect from a reference video is reused as visible content, it still belongs under `<Subject N>`. `<Video N>` identifies the asset or structural source and does not replace subject labels.
### 2.4 `<Audio N>`
`<Audio N>` represents a standalone audio asset or an enabled synchronized audio track from a reference video. Common uses include:
- Copying all or part of an audio signal
- Referencing a background-music style
- Referencing a speaker's voice timbre and delivery
- Using dialogue, lyrics, or sound effects from the original audio
- Referencing beat, rhythm, or audio continuity
When an `<Audio N>` explicitly corresponds to a target speaker, reuse that speaker's global ID in the definition: write `<Subject N> (Sx)` when the speaker maps to a defined subject, or use a stable voice description followed by `(Sx)` otherwise. The ID comes from the target video's global speaker order and is not independently assigned or renumbered in the audio definition. See Section 5.4 for the speaker-numbering rules:
```text
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).
```
When one audio asset serves multiple roles, describe those roles in one natural sentence rather than creating additional subsections.
### 2.5 Visual and Audio Tracks from the Same Reference Video
`<Video N>` and `<Audio N>` are numbered independently. Each index indicates only the label's order within its own category and does not encode a pairing between the two categories. The same reference video may therefore correspond to `<Video 1>` and `<Audio 2>`; different indices do not prevent them from coming from the same source asset.
An ordinary reference video does not create `<Audio N>` merely because the file contains sound.
An `<Audio N>` definition primarily states the audio's role and does not have to name the `<Video N>` it comes from. State the shared source only when needed to remove provenance ambiguity, for example:
```text
<Video 1> is the source video for the target video edit.
<Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.
```
## 3. `summary`
This section uses one short English paragraph to summarize the target video and its reference relationships. It begins with a square-bracketed task-type prefix:
```text
[reference generation] ...
[video editing + reference generation + audio reuse] ...
```
Choose task types according to the actual role each reference asset plays in the target video:
| Task type | When to use it |
| --- | --- |
| `keyframe completion` | An image serves as the target video's first frame, keyframe, last frame, edited keyframe, or another concrete frame anchor |
| `reference generation` | An image, video, or audio asset provides generation guidance for a character, scene, style, action, camera movement, storyboard, and so on, without serving as a concrete frame or as the source video being edited or continued |
| `video editing` | An existing source video is directly modified; editing an image or generating between still keyframes does not belong to this type |
| `video continuation` | New content continues, extends, resumes, or transitions from an existing source video |
| `audio reuse` | The same audio signal is reused in full or in part |
| `audio reference` | The audio signal is not copied directly; only its music style, timbre, dialogue or lyric content, sound-effect texture, beat, or continuity is referenced |
When a task satisfies multiple relationships, combine the task types with ` + ` and do not repeat a type. For example, continuing from a source video while using an image as the last frame is written as `[video continuation + keyframe completion]`. Editing a source video while retaining its original audio may be written as `[video editing + audio reuse]`.
The mere presence of video or audio does not automatically create a corresponding task type. If a reference video provides only camera movement, cuts, or rhythm, it normally belongs to `reference generation`. Use `video editing` or `video continuation` only when that video is directly edited or continued.
When editing a source video, use `audio reuse` as well if its original audio remains audible. When continuing a source video without directly copying the audio signal, use `audio reference` if the new audio only continues the original track's audible characteristics.
The summary uses the previously defined `<Subject N>`, `<Picture N>`, `<Video N>`, and `<Audio N>` labels to describe the main subjects, shot flow, and roles of the reference assets. Do not introduce new reference labels in this section.
For video-editing tasks, begin the summary after the task-type prefix with:
```text
The target video is an edited version of <Video 1>.
```
## 4. `retention_analysis`
This section describes how each piece of referenced content is preserved, transferred, copied, or referenced in the target video. Use one line for each reference label and preserve the meaning established in `subject_definitions`.
### 4.1 Visible Content
`<Subject N>`, `<Picture N>`, and `<Video N>` use the following relationship markers. These markers are fixed English values in the output format:
| Relationship marker | Meaning |
| --- | --- |
| `fully_preserved` | The defined role of the referenced content is fully preserved |
| `partially_preserved` | The referenced content is still used, but some defined characteristics are changed or only partially retained |
| `attribute_transfer` | Referenced characteristics are transferred to a different identifiable target subject |
| `weak_reference` | Only broad similarity in style, category, composition, or atmosphere is retained |
Subject entry:
```text
<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...
```
Picture entry:
```text
<Picture 2> ([Shot 1] first frame): fully_preserved - ...
```
Video-structure entry:
```text
<Video 1> (cut and pacing structure): weak_reference - ...
```
### 4.2 Audio
`<Audio N>` uses the following relationship markers:
| Relationship marker | Meaning |
| --- | --- |
| `fully_copy` | The complete source audio serves as the target video's complete final audio track |
| `partially_copy` | Only part of the timeline or selected audio layers are copied, or other sounds are added, removed, or replaced after copying |
| `reference` | The signal is not copied directly; only timbre, rhythm, music style, dialogue content, or sound texture is referenced |
| `weak_reference` | Only broad similarity in category or atmosphere is retained |
```text
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
```
```text
<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.
```
Choose each relationship marker only within the reference role already defined for that label in `subject_definitions`. Do not treat newly added actions, backgrounds, or plot events in the target video as losses of reference fidelity.
## 5. `detailed_description`
This is the main body of a full-reference rewrite. It describes visuals, actions, sound, and dialogue shot by shot in target-video playback order and inserts reference labels where they apply.
### 5.1 Basic Format
The basic format follows the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA):
- Write the body in English. Preserve the original language of dialogue, lyrics, and visible text.
- `[Shot 1]` marks the opening shot and has no timestamp. Later shots use `[Shot N] At MM:SS.mmm, ...` to mark cut times.
- Write camera movement as natural English within the current shot, including movement type, amplitude, and speed when they need to be expressed.
- Give vocal sources stable `(S1)`, `(S2)`, and subsequent IDs. Write dialogue and lyrics as `<d>[Language] ...</d>`.
- Use `<scenetrans>`, `<cutoff>`, and the corresponding continuity descriptions for dialogue crossing a cut, speech truncated by the video ending, and continuous audio across shots.
For complete rules and examples covering camera vocabulary, group speech, voice-over, dialogue across cuts, and visible text, see the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).
### 5.2 Full-Reference Mode Differences
| Dimension | T2VA | Full-reference mode |
| --- | --- | --- |
| Main field | `integrated_multimodal_description` | `detailed_description` |
| Style opening | Written after `[Shot 1]` | Established in one or two English sentences before `[Shot 1]` |
| Reference information | Does not use full-reference labels | Inserts `<Subject N>`, `<Picture N>`, `<Video N>`, and `<Audio N>` at their first appearance and where their roles apply |
| Audio relationships | Describes the target video's own sound | Cites `<Audio N>` in the corresponding shot or audio phase and states whether the signal is copied or referenced |
Opening example:
```text
The target video is in a cinematic, literary music-video style with soft lighting and a slightly desaturated color palette.
[Shot 1] The scene opens in a crowded urban street...
[Shot 2] At 00:09.000, the shot cuts to an extreme close-up...
```
For generation tasks, `detailed_description` is normally 350-500 English words. Dialogue-dense content prioritizes fitting the complete spoken timeline rather than mechanically reaching a word count. Video-editing descriptions scale with the complexity of the source video and do not have to follow the generation-task range. A single shot does not automatically justify a shorter description; distribute detail across multiple shots according to their information load.
### 5.3 Using Reference Labels in Shots
At the first clear appearance of an important `<Subject N>`, describe its referenced characteristics, position in the frame, and current action within what is actually visible in the shot. Continue using the same label in later shots without redefining what the label represents.
Use natural phrasing for concrete frame anchors:
```text
the shot begins from <Picture 1>
the shot's keyframe corresponds to <Picture 2>
the shot ends on <Picture 3>
```
When editing or continuing an original video, cite `<Video N>` naturally where its source state, structure, or continuation relationship applies. Cite `<Audio N>` in the shot or semantic phase where the audio relationship is active.
### 5.4 Speakers, Audio Sources, and Dialogue
The basic speaker-ID and `<d>` formats follow T2VA. When a referenced subject physically speaks, retain both the visual reference label and the speaker ID:
```text
<Subject 2> (S1) turns toward the woman and says, <d>[English] Last summer, I went to my grandfather's house. He talked about you.</d>
```
`<Subject N>` identifies the referenced subject, while `(Sx)` identifies the actual speaker. When the subject speaks, write `<Subject N> (Sx)`. If the same subject speaks off-screen, keep the same form and mark it as `off-screen`. When the speaker does not correspond to a defined subject, use a stable voice description followed by `(Sx)`.
When verbal content is only a cue within a directly reused BGM or complete soundtrack, and no person, character, narrator, or other independent vocal source physically produces it, use `<Audio N>` as the audible source and do not invent an additional `(Sx)`. If a concrete person, character, narrator, or other independent vocal source produces the voice, assign and reuse `(Sx)` for that source:
```text
When <Audio 1> reaches the phrase <d>[English] I'm lonely lonely lonely lonely lonely I'm lonely</d>, <Subject 1> performs the corresponding hand gesture without becoming a separate speaker source.
```
When dialogue, narration, or lyrics from reference audio are directly reused, or when the input prompt explicitly requests their reperformance, preserve the exact source words and original language inside `<d>`. Write `[unclear]` for unintelligible spans instead of guessing or paraphrasing them. Standardize punctuation to the basic written marks needed to express the sentence, such as `,`, `.`, `?`, and `!`; remove repeated tildes, emoji, bullets, and repeated or decorative punctuation. End complete statements, questions, and exclamations with `.`, `?`, or `!` respectively before `</d>`.
When only timbre, rhythm, emotion, or delivery is referenced, do not carry the original dialogue from the reference audio into the target video.
Assign `(Sx)` once according to the order of actual vocal events in the target video. Reuse the corresponding ID at every actual vocal event in `detailed_description`; an `<Audio N>` definition bound to a target speaker in `subject_definitions` also reuses the same `(Sx)` but never assigns a new one independently. Do not write `(Sx)` in `retention_analysis`. Verbal cues that exist only within a directly reused BGM or complete soundtrack use `<Audio N>`; voices physically produced by a concrete person, character, narrator, or other independent vocal source use `(Sx)`.
## 6. `overall_soundscape` and `non_diegetic_music`
The definitions of these two sound categories follow the Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA).
`overall_soundscape` summarizes ambience and physical sounds across the full video. Dialogue, singing, and sound events synchronized to a particular shot remain in `detailed_description`:
```text
overall_soundscape: Quiet indoor room tone and a low ventilation hum continue throughout the video.
```
`non_diegetic_music` describes background music that the characters cannot hear and that is audible only to the audience. When music is present, state its instrumentation, tempo, and dynamic development:
```text
non_diegetic_music: A restrained solo-piano score at a slow tempo, with sustained low cello underneath and no swell.
```
When reference audio is used, state its copy or reference relationship only in the section that matches the audible layer: ambience and sound effects belong in `overall_soundscape`, while audience-only score belongs in `non_diegetic_music`. If the same audio provides both kinds of content, describe the corresponding relationship in each section:
```text
overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.
non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.
```
Write complete dialogue and lyrics only inside `<d>` in `detailed_description`; do not repeat them in these two sections.
## 7. Complete Example
<details>
<summary>Show the complete example</summary>
```text
subject_definitions:
<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.
<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.
<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.
<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.
<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.
summary:
[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.
retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.
<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.
<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.
<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.
detailed_description:
The target video uses a realistic multi-camera sitcom style with warm indoor lighting.
[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.
[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.
[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.
overall_soundscape:
Soft indoor coffee-shop room tone continues throughout the scene.
non_diegetic_music:
N/A
```
</details>
以“黑蚂蚁在枯叶上迎战暴雨水滴” 为例,手把手带你从零写出一份 100% 符合官方标准规范的中文 I2V(图生视频)提示词
📌 设定背景与准备素材
- 生成任务:图生视频 (I2V / I2VA)
- 输入素材:已生成好的首帧图
<Picture 1>(画面为:一只光泽黑工蚁趴在枯黄树叶上,背景是湿润泥土)
🛠️ 手把手编写 4 步法
第一步:写【首帧对齐指令】(Part One)
在提示词的第一行,必须声明 <Picture 1> 在视频 0.00 秒时作为首帧完全对齐。空一行后再写后续内容。
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.编写要领:这一行是写给模型解析器看的固定指令,保持英文原文即可,用于“锁定”首帧。
第二步:写【核心画面与动作】(integrated_multimodal_description:)
这是提示词的主体,采用标准公式:
[Shot 标记] + [继承首帧画风/主体] + [动态发展] + [自然运镜] + [说话人与台词]。
1. 镜头标记与画风继承:
写出 [Shot 1](注意: Shot 1 绝对不带时间戳)。说明继承 <Picture 1> 中蚂蚁的姿态与场景。[Shot 1] 实拍、电影感微观摄影。画面继承 <Picture 1> 中枯叶上的光泽黑工蚁姿态、枯叶纹理及暗色背景氛围。
2. 画面动态与具象动作:
描述 0 秒后的动态演变,用具体的物理动作代替抽象词。在 0.00 秒时,一滴巨型水珠从上方急速坠落,以超慢动作砸在枯叶紧挨着的湿泥上,剧烈冲击力导致泥水与水花向四周飞溅。水滴爆裂瞬间,蚂蚁迅速压低身体,用脚紧紧抓牢枯叶边缘以抵御冲击波。
3. 嵌入自然运镜:
将运镜的“类型 + 幅度 + 速度”作为自然句子写入,不要独立堆砌标签。镜头以慢速、小幅度向蚂蚁头部推近。
4. 说话人与台词格式:
指定说话人 ID (S1),台词必须放在 <d>[Chinese] ...</d> 中,保持原字原句。深沉稳重的电影感男旁白(S1)说道:<d>[Chinese] 暴雨倾盆而下,每一滴都是致命的考验。</d>
合并第二步的完整段落:
integrated_multimodal_description: [Shot 1] 实拍、电影感微观摄影。画面继承 <Picture 1> 中枯叶上的光泽黑工蚁姿态、枯叶纹理及暗色背景氛围。在 0.00 秒时,一滴巨型水珠从上方急速坠落,以超慢动作砸在枯叶紧挨着的湿泥上,剧烈冲击力导致泥水与水花向四周飞溅。水滴爆裂瞬间,蚂蚁迅速压低身体,用脚紧紧抓牢枯叶边缘以抵御冲击波。镜头以慢速、小幅度向蚂蚁头部推近。深沉稳重的电影感男旁白(S1)说道:<d>[Chinese] 暴雨倾盆而下,每一滴都是致命的考验。</d>第三步:写【环境与物理音效】(overall_soundscape:)
总结整段视频中的环境音、物理碰撞音(1-4 句)
绝对不在此处重复台词、人声或背景音乐
overall_soundscape: 持续的微观雨声环境音。水滴砸入泥土时发出沉闷而巨大的撞击轰鸣声,紧接着是液体飞溅和湿泥挤压的声音。第四步:写【背景配乐】(non_diegetic_music:)
non_diegetic_music: 开场为慢节奏的低音大提琴长音,在水滴砸落瞬间加入沉重的管弦乐重击与低音弦乐渐强,营造出紧张的氛围。📋 最终完整的官方标准提示词
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] 实拍、电影感微观摄影。画面继承 <Picture 1> 中枯叶上的光泽黑工蚁姿态、枯叶纹理及暗色背景氛围。在 0.00 秒时,一滴巨型水珠从上方急速坠落,以超慢动作砸在枯叶紧挨着的湿泥上,剧烈冲击力导致泥水与水花向四周飞溅。水滴爆裂瞬间,蚂蚁迅速压低身体,用脚紧紧抓牢枯叶边缘以抵御冲击波。镜头以慢速、小幅度向蚂蚁头部推近。深沉稳重的电影感男旁白(S1)说道:<d>[Chinese] 暴雨倾盆而下,每一滴都是致命的考验。</d>
overall_soundscape: 持续的微观雨声环境音。水滴砸入泥土时发出沉闷而巨大的撞击轰鸣声,紧接着是液体飞溅和湿泥挤压的声音。
non_diegetic_music: 开场为慢节奏的低音大提琴长音,在水滴砸落瞬间加入沉重的管弦乐重击与低音弦乐渐强,营造出紧张的氛围。