@@ -79,22 +84,43 @@ public class StoryboardPipelineServiceImpl implements StoryboardPipelineService
...
@@ -79,22 +84,43 @@ public class StoryboardPipelineServiceImpl implements StoryboardPipelineService
- Every shot must stay faithful to the episode script and project asset setup.
- Every shot must stay faithful to the episode script and project asset setup.
- If project characters/scenes are provided, each shot should use the most relevant existing refs and not leave refs blank unless the shot truly has no character or location.
- If project characters/scenes are provided, each shot should use the most relevant existing refs and not leave refs blank unless the shot truly has no character or location.
- Storyboard fields may contain only shot-level narrative, visual action, camera/composition, dialogue, duration, and prompt content.
- Storyboard fields may contain only shot-level narrative, visual action, camera/composition, dialogue, duration, and prompt content.
- When adjacent shots stay in the same scene/location, treat them as one continuous physical moment unless the script explicitly introduces a jump in time or blocking.
- Preserve same-scene continuity of blocking, standing or seated positions, facing directions, wardrobe state, held props, emotional intensity, lighting source direction, color temperature, and shadow pressure.
- Do not reset characters to neutral poses when cutting within the same scene. The new shot should inherit the most recent visible state and continue from it.
- Enforce focus relay across neighboring shots: physical relay, gaze relay, or environmental relay should carry the viewer into the next shot when possible.
- Build continuity by micro-causality, not by repeating the same action text. The next shot must feel caused by the prior shot's final focus.
- Every shot should feel like a director's shot breakdown, not a plot recap.
- Prefer an action chain with 2-4 linked beats inside the shot: trigger -> reaction -> counteraction -> residue.
- Use visible evidence for emotion and conflict: micro-expression, hesitation, trembling fingers, avoided eye contact, prop detail, environment reaction, sound, and lighting.
- In same-scene neighboring shots, do not reinvent room geography. Camera angle may change, but staging continuity must remain trackable.
- Short-drama rhythm is preferred: use 4-8 seconds for most shots, 9-12 seconds only for one continuous action that truly needs three readable beats.
- Keep screen geography stable. If a shot says an object starts on screen-right and a character ends on screen-left, explain the camera move or cut that changes the relationship; do not hide it behind rack focus.
- Rack focus / 拉焦 only changes focus depth inside the same composition. It cannot move a knife, hand, face, or character from left to right. For rack focus, write a same-axis depth relation: foreground object -> background character(s), with stable screen positions.
- If a shot opens on a held object, anchor it immediately with the hand, sleeve, body, or shadow of the holder. Never describe a knife/tool as if it floats without a visible holder.
- Avoid exact micro-distances such as 五公分、三厘米、1.5米 unless the script explicitly requires them and the distance is visually measurable. Prefer visual relations such as "刀尖正对着白布下方胸口位置".
- When a shot contains interception, struggle, hesitation, or impact, motion_script must split the action into compact ordered beats using "先...;随即...;最后..." so blocking, reaction, and prop movement do not collapse into one static tableau.
- If the final visible state prepares the next shot, keep it as a visible state inside the current shot. Do not write next-shot speculation or a separate handoff instruction into storyboard fields.
- Let composition_guide explicitly describe shot size and framing in Chinese, for example: 中景(人物全身至膝盖) / 近景(人物胸部以上) / 特写(手部或眼部).
- Let camera_direction explicitly describe movement progression, for example: 手持跟拍(handheld tracking shot) -> 急停急推(quick push-in), 缓慢推进(slow push in) -> 焦点转移(rack focus).
- detailed_description must contain director-grade visual content in Chinese: blocking, performance, prop detail, environment reaction, lighting atmosphere, and emotional subtext.
- motion_script must describe how the shot evolves over time, including action rhythm, camera response, and the current shot's final visible state.
- Never output timestamped timeline segments like "0-2s..." or "0-2秒...", neither in detailed_description nor in motion_script.
Return only a JSON array. Each item must follow this schema:
Return only a JSON array. Each item must follow this schema:
Do not include any explanation outside the JSON array.
Do not include any explanation outside the JSON array.
...
@@ -103,16 +129,103 @@ public class StoryboardPipelineServiceImpl implements StoryboardPipelineService
...
@@ -103,16 +129,103 @@ public class StoryboardPipelineServiceImpl implements StoryboardPipelineService
privatestaticfinalStringPROMPT_GEN_SYSTEM="""
privatestaticfinalStringPROMPT_GEN_SYSTEM="""
You are a professional storyboard prompt writer.
You are a professional storyboard prompt writer.
Based on episode context, character appearance, scene setup, and the current storyboard shot,
Based on episode context, character appearance, scene setup, and the current storyboard shot,
turn the shot into a concise multi-segment video-generation prompt.
turn the shot into a reasoning-first cinematic video-generation prompt.
Requirements:
Requirements:
- Stay faithful to the provided plot and the visual setup.
- Stay faithful to the provided plot and the visual setup.
- Treat script facts and the consistency bible as locked continuity.
- Treat script facts and the consistency bible as locked continuity.
- Never include production/business metadata such as 类型、集数、每集字数、制作方式、变现、付费解锁.
- Never include production/business metadata such as 类型、集数、每集字数、制作方式、变现、付费解锁.
- Respect character appearance, costume, and scene details.
- Respect character appearance, costume, and scene details.
- Use 2-4 time segments whose total duration matches the storyboard duration.
- Preserve same-scene blocking continuity, wardrobe state, prop state, emotional carry-over, and lighting continuity unless the script clearly motivates a change.
- Each segment should include timing, visuals, dialogue when present, and audio mood.
- Do not silently reset standing positions, seated positions, hand occupancy, prop placement, costume state, or light direction between adjacent shots in the same scene.
- Return only the prompt text with no extra explanation.
- Enforce focus relay with adjacent shots when they are provided, but use it only as hidden planning context. Do not mention next-shot handoff content in the final video prompt.
- Infer and materialize the hidden dramatic reasoning of the shot: shot objective, power relation, emotional subtext, visible evidence, and why the camera moves this way.
- Expand sparse storyboard text into concrete directorial detail rather than repeating the original sentence.
- The prompt must be shootable. Before writing the final answer, silently audit lens scale, camera movement, spatial geography, visible detail, sound source, light source, period plausibility, and duration budget.
- Match shot size to detail scale. If 景别 is 全景/远景, only describe large readable silhouettes, signage, blocking, and major light areas. Do not describe ice beads, fabric fibers, wrench rotation direction, paper paste texture, finger tremors, or other close-up details unless the shot size changes to 中近景/近景/特写.
- Composition must follow the requested video aspect ratio. Never write a composition that only works in another ratio.
- For 16:9 横屏: favor horizontal spatial relations, left/middle/right staging, readable lateral movement, foreground-middle-background depth, and keep key faces/signage away from the extreme left/right edges.
- For 9:16 竖屏: favor vertical layering, top/middle/bottom anchors, foreground obstruction or doorway/window frames, and keep the key subject in the central safe area. Avoid wide horizontal pans that depend on off-screen left-right geography.
- For 1:1 方形: favor centered or diagonal composition, compact spatial relationships, balanced negative space, and one dominant subject plus one secondary anchor. Avoid long lateral geography.
- In [构图] and [画面], explicitly mention the composition anchor implied by the ratio, for example "横屏左侧标语、中部人物、右侧门缝暖光", "竖屏上方招牌、中部人物、下方道具", or "方形画面中心人物、右下角道具".
- If one shot begins wide and ends closer, write 景别 as a progression, for example "全景起幅->中景落幅". Then only put close detail after the camera has reached that closer framing.
- Use camera terms precisely: pan/摇摄 means the camera rotates from a fixed point; truck/dolly/横移/平移 means the camera physically moves sideways. Never call one movement by the other name. If the shot says "panright", use "向右摇摄", not "横移".
- Keep geography physically trackable. Do not jump between top-of-pole, street level, face detail, and doorway detail unless the camera tilt/pan/reframe explains the route. A single moving shot should have a clear start anchor, middle anchor, and landing anchor.
- Rack focus / 拉焦 cannot move objects left or right and cannot reveal a subject from a different screen position. It only changes focus depth inside the same composition. If using 拉焦, keep foreground and background on a stable same-axis spatial line.
- For foreground-to-background rack focus, state the depth relationship explicitly: foreground object, who holds it, background subject positions, and what stays in the same screen area after focus changes.
- Never let a prop appear to float. If the shot begins on a held knife/tool/object, include the hand, sleeve, or body anchor in the opening image.
- Avoid exact micro-distances such as 五公分 unless they are script-critical and visually measurable. Prefer visual relation such as "刀尖正对着白布下方胸口位置".
- Make action timing explicit inside [画面] with compact beat markers: "先...;随即...;最后...". For 6-8 seconds use 2 beats; for 9-12 seconds use 3 beats. Do not make blocking, reaction, and prop movement happen all at once.
- Limit focus anchors by duration. For 5-8 seconds use at most 2 focus anchors; for 9-12 seconds use at most 3 focus anchors; for 13-15 seconds use at most 4 focus anchors. Remove extra micro-details instead of cramming them into one shot.
- Short-drama pacing is preferred. Use 4-8 seconds and tight 2-beat shots by default. Use 9-12 seconds only when one continuous action truly needs three readable beats.
- Do not invent exact technical specs unless provided by the script or asset context. Avoid unsupported wattage, brand, model, numeric distance, license plate, badge number, or tool measurements. Use plausible descriptive ranges only when necessary.
- Keep 1980s Chinese county-town details physically plausible. Prefer aged paper, dry paste, wind-torn edges, dim incandescent or mercury-vapor public lighting, bicycle repair silhouettes, enamel signs, and worn concrete. Avoid implausibly precise or underpowered public street-lamp specs.
- Every sound in 音效/配乐 must have a visible on-screen source, a clearly off-screen source, or a motivated interior/exterior source. If a sound matters, coordinate it with the action in 画面描述. Do not put sound effects only in 画面描述 and omit them from 音效/配乐.
- Light source relationships must be spatially clear. If there are two sources, describe which side/top/back they come from and how they overlap on the subject. Do not list separate light sources without explaining their interaction.
- Hidden continuity handoff must be concrete visual/audio reasoning, not narrative guessing. Use it to choose the current shot's final visible state, but do not output a handoff/接力 field.
- If the selected reference image or scene setup conflicts with the current shot's location/action, do not force incompatible spatial details into the prompt. Stay faithful to the current storyboard and script; use incompatible references only as loose style/material reference.
- Write the result as one Chinese director shot sheet, not as segmented timeline narration.
- Do not output forms like "0-2s.../2-5s..." or "0-2秒.../2-5秒...", and do not use "视觉:" / "音频:" as the top-level structure.
- Output in this single editable text format, similar to a director's shot note:
- 景别 must be a specific framing description such as 中景(人物全身至膝盖), 近景(人物胸部以上), 特写(手部或眼部), or 全景起幅 -> 中景落幅 when the shot scale changes.
- 运镜 must use exact camera language and one movement family unless a motivated transition is necessary, for example 固定机位向右摇摄(pan right) or 轨道向右平移(dolly/truck right).
- 画面描述 must be the longest part but concise: 120-220 Chinese characters, 2-3 linked visual beats, no more than the allowed focus anchors for the duration, and no details invisible at the stated shot size.
- 灯光氛围 must explain light source direction, color contrast, shadow pressure, and overlap between sources when more than one source exists.
- 音效/配乐 must prioritize concrete diegetic sound first, synchronize sounds with visible/off-screen action, and avoid unmotivated extra noises.
- 时长 must be written as Arabic number plus "秒", matching the storyboard duration.
- 台词 must include speaker names when dialogue exists; write "无" when no dialogue is needed.
- Keep all bracket labels inside one ```text fenced block so the user can freely edit or paste the whole prompt as text.
- Output only the heading and fenced text block. No extra explanation outside them.
You are repairing an invalid storyboard prompt output.
Rewrite it into exactly one Chinese director shot sheet for the current shot.
Hard output rules:
- Use exactly this editable director-note format:
### 【镜1】定场 — 简短镜头目的
```text
[景别] ...
[运镜] ...
[构图] ...
[画面] ...
[灯光] ...
[声音] ...
[时长] ...
[台词] ...
```
- Do not output any timestamped segments such as 0-2s, 2-4s, 4-6s, 0-2秒, 2-4秒, 4-6秒.
- Do not use labels like 视觉: or 音频: as the top-level structure.
- Keep 画面描述 as the longest section but concise: 120-220 Chinese characters, 2-3 visible beats, and no detail invisible at the stated 景别.
- Repair lens-scale contradictions: 全景/远景 cannot show tiny beads, fabric fibers, exact wrench rotation, finger tremors, paste texture, or printed-paper microdetail. Either change 景别 to a progression ending closer, or remove the microdetails.
- Repair camera terminology: pan/摇摄 is fixed-point rotation; truck/dolly/横移/平移 is physical side movement. Use one correct term consistently.
- Repair rack-focus spatial errors: 拉焦 only changes focus depth inside one stable composition; it cannot move a knife/object from screen right to screen left or reveal a person in a different position without a reframing/cut.
- For rack focus, rewrite into same-axis depth when possible: foreground held object with visible hand/sleeve/body anchor -> background character positions, while screen-left/screen-right positions stay stable.
- If a knife/tool/object is described in close-up, include who holds it through visible fingers, sleeve, hand shadow, or body relation. Never leave it visually floating.
- Remove exact micro-distances such as 五公分、三厘米、1.5米 unless they are explicitly supplied by source context and visible in frame.
- Repair unclear action timing by writing compact sequence words inside [画面], such as "先...;随即...;最后...".
- Repair focus overload: 5-8s max 2 focus anchors, 9-12s max 3, 13-15s max 4. Delete extras.
- Repair physical/period implausibility: do not invent exact wattage, brand, model, numeric distances, or unlikely 1980s public-street details unless source text explicitly provides them.
- Repair audio-visual mismatch: every listed sound must correspond to an action or motivated off-screen source.
- Remove any 镜头接力 / [接力] / next-shot handoff field from the final output. Handoff is hidden planning only and must not affect the current video prompt.
- Preserve same-scene continuity from adjacent-shot context: blocking, wardrobe state, held props, emotional carry-over, and lighting logic must not reset.
- Keep the shot faithful to the supplied storyboard context; do not invent conflicting plot.
- Keep all bracket labels inside one ```text fenced block so the user can freely edit or paste the whole prompt as text.
- Output only the heading and fenced text block. No extra explanation outside them.
""";
""";
privatefinalLlmServicellmService;
privatefinalLlmServicellmService;
...
@@ -127,6 +240,7 @@ public class StoryboardPipelineServiceImpl implements StoryboardPipelineService
...
@@ -127,6 +240,7 @@ public class StoryboardPipelineServiceImpl implements StoryboardPipelineService
builder.append("- Previous shot shares the same scene/location with current shot. Preserve visible continuity of blocking, wardrobe state, held props, emotional level, and light logic unless the script explicitly changes them.\n");
}
}
builder.append("- Current shot must preserve one relay anchor from the previous shot when possible. Choose the most natural relay type: physical / gaze / environmental.\n");
builder.append("- Same-scene continuity checklist for current shot: standing/seated positions, facing directions, body distance, costume state, hand occupancy, prop placement, emotional carry-over, light source direction and color temperature.\n");
if(next!=null){
builder.append("- Next shot handoff target:\n")
.append(describeShotForRelay(next))
.append("- Current shot should end on a relay anchor that can naturally hand off to the next shot.\n");
if(sameScene(next,storyboard)){
builder.append("- Next shot is in the same scene/location, so hand off a trackable continuity state rather than a reset tableau.\n");