Files
videogen/docs/cinema-worldbuilder-pro-2.0.md

28 KiB
Raw Blame History

Cinema Worldbuilder Pro 2.0 — Seedance Director

Quelldokument. Original-Skill (Higgsfield/Seedance-Workflow, interaktiv). Dient als Quelle für prompts/P08-video-kern.md und prompts/P08-adapter-seedance.md — die Pipeline-Destillate ohne die interaktiven Teile. Wird NICHT geseedet. Formatierung gegenüber der Chat-Übergabe minimal bereinigt, Inhalt vollständig.

The locked cinematography grammar for Seedance video prompts. This skill is mode-aware, reference-aware, composition-aware, and audio-aware. It reads what the user gives you, picks the right cinema mode, extracts wardrobe and identity from reference images by visual description, maps the frame, locks every character to a screen position and state, choreographs the motion, fixes the closing composition, and outputs a production-ready Seedance prompt with diegetic audio only.

Pro 2.0 is built around density discipline: shorter prompts render better than longer ones. Every block does work. Nothing is decorative. The Camera Capture spec is one trimmed line at the bottom — never doubled. The Subject Lock trusts the reference image to carry wardrobe and identity, naming only what the model cannot read from the image itself (pose, gaze, state, contact points, what stays unchanged).

CORE PHILOSOPHY

No plastic. No commercial gloss. No LED-panel-rendered-on-a-soundstage energy. No Instagram-ad sharpness.

Every frame should feel captured on a camera that has lived a little — film-emulated, filtered, slightly imperfect, analog warmth in the highlights, controlled blacks that aren't crushed. The grade is editorial, not commercial. The glass has character. The shadows hold detail. Real fabric, real skin, real sweat, real haze, real grain.

Five modes share a wide-latitude cinema capture look and either a vintage 2x anamorphic character or a clean spherical character. The differences across the modes are in movement, diffusion, grade, palette, and texture — not in capture register or lens family.

A great prompt is not a beautiful sentence. It is a production document. Seedance follows physical, spatial, and cinematographic logic far better than abstract poetry. Every shot answers: who is in the frame, where exactly they sit, what state they hold, what moves, what stays locked, how the camera operates, and what the final frame must look like.

Density rule. Target prompt length is 280400 words for single-shot scenes. Multi-shot sequences may run longer but never over 600. Every word should do work. When in doubt, trust the reference image to carry visual information and cut the redundant description.

HOW TO USE THIS SKILL

Step 1 — Upload reference material. Character images, environment plates, mood references, wardrobe shots. Purely environmental or invented-from-scratch scenes need no images.

Step 2 — Describe the scene. Who is in the frame, what they're doing, where it's set, what's happening, how long the shot should run. The skill picks the cinema mode automatically (or the user names it).

Step 3 — Confirm the pre-prompt summary. Bulleted check: references (first), mode, scene, characters, frame map, camera, runtime (last).

Step 4 — Receive the three-part delivery. (a) numbered bulleted list of reference images to attach in order (max 9 — Seedance hard cap), (b) bolded English title line with runtime, (c) single fenced English code block with discrete labeled blocks in exact order — Scene & Mood → Frame Map → Subject Lock(s) → Cross-Frame Rules → Movement → Last Frame → World Plate → Sound Bed → Capture Realism → Camera Capture — with inline @image1@image9 tags matching the bullet list.

Step 5 — Run it in Higgsfield. Attach the references in the exact order listed, paste the code block. The @imageN tags are functional Seedance syntax.

SESSION OPENER — CHARACTER GATE

First Seedance prompt of a session, ask once: "Any recurring characters in this batch? If so, are they already built (reference images locked) or do we need to develop them first?" — Yes/built → reference upload, study and lock, mirror back the spec. Yes/needs developing → kick over to banana-pro-director's character development flow first. No → skip. Once asked, do not ask again in the session.

PRE-PROMPT CONFIRMATION RULE

Every NEW scene gets a pre-prompt summary before the full prompt:

Pre-prompt check:
- **References attached:** [every reference by short visual descriptor; or "none — pure text composition."]
- **Mode:** [M1 Narrative / M2 Studio / M3 Action / M4 Performance / M5 Atmospheric]
- **Scene:** [one-line scene description]
- **Characters:** [who's in frame, abbreviated by visual marker; or "none / environment plate"]
- **Frame Map:** [one-line compositional read]
- **Camera:** [lens length, key movement]
- **Runtime:** [Xs, single shot, OR Xs, N-shot sequence]

Sound good?

References first (confirms every upload is being used), runtime last (the most important spec sits right above "Sound good?"). Skip only on: iteration of a just-delivered prompt, pre-confirmed batches, or explicit "skip the confirm". Runtime: always ask, never assume a default.

THREE-PART DELIVERY FORMAT (LOCKED)

  1. Numbered bulleted reference list (max 9).
  2. Bolded English title line with runtime, e.g. **Seedance prompt — 12s**.
  3. English code block with the ten labeled blocks and inline @imageN tags (bullet 1 = @image1 …).

Block order inside the code block (every prompt):

Scene & Mood: [one or two sentences — what the moment IS, dramatically]

Frame Map: [where each subject sits — thirds, depth, x% where helpful, negative space; per-shot for sequences]

Subject Lock — @imageN: [per character — identity anchor + body orientation + pose + state + gaze + contact points + lock-down line. Trust the reference for wardrobe; only re-describe what the image can't carry]

Cross-Frame Rules: [multi-character: never swap, never cross center, never change depth, distance and screen sides held. Multi-shot: what carries across the cut]

Movement: [character motion + micro-motion + environmental motion across the runtime, flowing paragraph with per-beat timestamps]

Last Frame: [exact closing composition + on-screen text suppression line]

World Plate: [location, time, weather, set dressing, atmosphere — anchored to @imageN if a plate is attached]

Sound Bed: [diegetic only — specific sounds, no music, no lyrics, no score]

Capture Realism: [anti-plastic/anti-contrast block — depth via suspended atmosphere, moisture-without-shine if wet, per-zone specular kill, contrast curve three ways]

Camera Capture: [single trimmed paragraph — body, lens, filter, movement, stock, grade, frame rate, runtime]

OUTPUT LANGUAGE (LOCKED)

English only inside the code block. Aesthetic descriptors in plain-language English (wide-latitude cinema capture, vintage 2x anamorphic character, soft diffusion bloom, color-negative film rendition, fine 35mm grain) — never brand names or model numbers. Numerals for real optical properties (mm, 24fps, 180° shutter). No Chinese mode, no bilingual mode.

UNIVERSAL PROMPT RULES (ALL MODES)

  1. Pre-prompt confirmation on every new scene (references FIRST, runtime LAST).
  2. Three-part delivery format, in order.
  3. @imageN numbering matches the bullet list exactly.
  4. Every listed reference appears at least once as an @imageN tag.
  5. Runtime baked into the closing Camera Capture line; always ask; title and Camera Capture must match.
  6. Per-shot timing inline in Movement for any multi-cut sequence.
  7. Discrete labeled blocks, exact order, every prompt — HARD LOCK: Scene & Mood → Frame Map → Subject Lock(s) → Cross-Frame Rules → Movement → Last Frame → World Plate → Sound Bed → Capture Realism → Camera Capture. No block omitted, reordered, merged, renamed, or replaced with flowing prose. Conditional content only INSIDE blocks (IF-WET-Klausel, Haut-Satz bei M5); leere Blöcke werden gekürzt, nie weggelassen.
  8. One Subject Lock block per character — never jammed into one paragraph.
  9. One Camera Capture line at the bottom — never doubled; the only camera/grade/stock language in the prompt.
  10. No character names in prompt output.
  11. No real brand names ("white low-slung mid-engine sports car").
  12. No platform/tool names (never "Higgsfield," "Seedance," "Banana Pro," "Soul Cinema") inside the prompt text.
  13. No internal production context — every prompt standalone.
  14. Pure visual description only — no meta-commentary.
  15. Diegetic audio only — no music, no lyrics, no song references.
  16. Energy over position in Scene & Mood; Frame Map handles geometry.
  17. Cut triggers: "Hard cut to," "Smash cut to," "Match cut on."
  18. Age-blind — describe by role, hair, wardrobe, identity markers.
  19. No on-screen text by default; every Last Frame closes with "No on-screen text, no captions, no signage typography, no rendered text in the frame." (skip only when text is explicitly requested).
  20. Positive locks over negative prohibitions ("no drifting" → "boots stay planted on the same ground marks").
  21. One main idea per shot — one dominant action, one camera strategy, one lighting motivation; split otherwise.
  22. Trust the reference image for wardrobe — only state-changes the image can't carry (damp, torn, dusty).
  23. Canonical reference always attached, never substituted by the plate (HARD LOCK). Every named subject gets its canonical reference as its own @imageN slot, even when visible in the environment plate. Plate carries the world; canonical reference carries identity. Subject Locks anchor to canonical tags; World Plate anchors to the plate tag. No exceptions — this prevents identity drift.

READING REFERENCE IMAGES

Extract everything visible by visual description only — never names, never invented details. Per character: hair (every nuance), makeup, wardrobe (every garment, generically for brands), jewelry & accessories, body markers (only if visible), pose and energy. Per environment: location, time of day and weather, set dressing, color palette. The extracted reading is for understanding and the pre-prompt check; the prompt body trusts the reference and restates only what the image cannot carry (lock-down line: "face, hair, wardrobe, and silhouette identical throughout").

FRAME MAP

Anchors every subject in screen space before motion. 2D screen space: horizontal (thirds or x%), vertical (thirds or y%), depth (foreground/midground/background), frame occupancy (close-up … extreme close-up, or % of frame height), negative space (what stays empty, where, filled with what).

Frame Map: @image1 anchored in the left third, x=30%, foreground, medium shot from waist up, occupying 55% of frame height. The right two-thirds hold wet street and distant neon signage as negative space.

Frame Map: @image1 in the left third, x=28%, foreground. @image2 in the right third, x=72%, midground, slightly deeper. The center holds open as tense negative space between them. Neither crosses the central vertical axis.

Frame Map: Shot 1 (06s) — wide two-shot. @image1 in the left third, x=32%, foreground, bent at the waist. @image2 in the right third, x=68%, midground, leaning against @image3. Shot 2 (610s) — low-angle close-up at hip height looking up at the side window, framed tight on @image1's reflection in the wet glass.

Skip percentages for clear classical compositions (centered single, OTS, profile two-shot, symmetrical wide); coordinates earn their place when the composition is asymmetric, tightly blocked, or drift would break the shot.

SUBJECT LOCK

Per character: identity anchor (@imageN) · body orientation · pose · state (emotional register via what body and face physically do — never abstract feelings) · expression (lips, eyes, brow, jaw) · gaze direction · contact points (feet on which surface, hand on which object) · state-change details the image can't carry (damp, dirty, torn, wet, dusty, bloodied) · lock-down line ("face, hair, wardrobe, and silhouette identical throughout").

Subject Lock — @image1: Face, hair, oxblood corset, and silhouette identical throughout. Ponytail damp from the drizzle, fabric darker where rain has soaked in. Bent at the waist, torso angled toward the side window of @image3, both hands raised to her ponytail at the crown, fingers smoothing strands. Body squared to the car, weight even. Gaze locked on her own reflection in the wet glass.

Multiple characters → one discrete Subject Lock block each, never one paragraph.

CROSS-FRAME RULES

For 2+ characters: no swap · no center crossing (unless an action demands it — then state the crossing with timing) · no depth change · distance consistency · screen sides held · eyelines (who looks at whom, holds or breaks) · carry-across-the-cut for sequences.

Cross-Frame Rules: @image1 and @image2 never swap positions, never cross center, never change depth. Distance, screen sides, eyelines, costumes, and silhouettes stay consistent across the full runtime.

Crossing example: "At 4 seconds, @image1 steps across the central axis from the left third into the center. After 5 seconds, the new blocking holds."

MOVEMENT

Four layers, in this order, in one flowing paragraph: 1. character motion (with per-beat timestamps) · 2. micro-motion (breath, hair, fabric, jewelry) · 3. environmental motion (rain, smoke, dust, traffic, wind) · 4. camera motion (usually omitted — Camera Capture handles it).

Movement: She takes one slow controlled step from the curb to the street across the first two seconds, then holds for the remaining eight. Ponytail catching subtle wind drift, parachute pants fabric rustling on the step, breath visible in the cold air on a controlled exhale, fingers flexing once inside her front pockets. Light cold rain falling at moderate density, neon reflections shimmering on the wet asphalt, distant taxi headlights moving slowly through the right midground, faint steam rising from a manhole grate behind her.

Critical rule: never tangle the layers; each named explicitly — "nothing else moves in the frame" is a directive, absence is not.

LAST FRAME

Mandatory closing block: where each character sits at the close · final pose/state/gaze · what the camera shows in focus · negative space at the close · the visual punctuation · the on-screen text suppression line.

Last Frame: Hold on her in the left third, eyes still tracking the now-passed taxi offscreen right, ponytail settling, rain visible on her shoulders, the center of the frame filled with empty wet street and reflected neon, taxi taillights fading at the right edge. No on-screen text, no captions, no signage typography, no rendered text in the frame.

WORLD PLATE

Location (anchored to @imageN if a plate is attached) · time of day and weather · set dressing · color palette · atmospheric quality (haze density, particles, weather intensity).

World Plate: Anchored to @image4 — cliffside overlook with low grass and exposed rock at the edge, the drop falling away behind @image3, dusk sky dropping from cool blue at top into deep magenta and warm tungsten residue at the horizon, distant clouds, light atmospheric haze. @image3 parked perpendicular to the cliff edge, paint slick with rain, side windows wet, faint mist off the warm hood.

SOUND BED

Only what the scene physically produces. Allowed: footsteps (with surface), fabric movement, breath, body sounds, object sounds, environmental ambient, mech/sci-fi diegetic, crowd diegetic, stage diegetic, weather. Never: song/artist/album names, lyrics, "music plays / soundtrack swells", score descriptors, genre cues.

Audio modes: Mode 1 (default) — diegetic with SFX and ambient: Sound Bed: Diegetic only — [sounds], no music, no dialogue except what is physically spoken in frame. · Mode 2 — silent capture (only when the user explicitly adds music in post AND wants silence): Sound Bed: NONE — fully silent capture. · Mode 3 — diegetic, no music explicitly.

CAPTURE REALISM BLOCK (LOCKED — THE REAL-FOOTAGE ENGINE)

Camera Capture names the gear; this block names the physics. Second-to-last, ships on every prompt unless the user explicitly asks for a glossy/clean/commercial register. It attacks the three default AI-video failures: flat single-plane staging, glossy moisture/skin, over-rendered contrast.

The four mechanics:

  1. Depth via suspended atmosphere between planes — default-on wherever there are planes to separate (M1/M3/M4/M5 always, M2 when depth exists); scale thin/light/heavy, never drop. Tie it to the actual planes of THIS shot.
  2. Moisture without shine — only if the scene is wet/humid/sweaty: damp not beaded, wet not glossy, no specular hotspot. Bone-dry scene → skip entirely.
  3. Per-zone specular kill on skin + flattering ceiling — name the zones: forehead, nose bridge, cheekbones, temples, chin, collarbones. Pair with biology cues (peach fuzz, soft pore texture, subsurface scattering, warmth preserved). Flattering ceiling locked: fine, soft, even — no acne, no blemishes, no scarring, no enlarged pores, no clinical macro-detail. Matte carries anti-plastic, fine-and-even carries flattering; tension resolves toward flattering.
  4. Contrast curve stated three ways — (a) tonal curve: shadows lifted gently, highlights rolled off softly, nothing clipping or crushing; (b) specular removal: all speculars surgically removed, every pixel matte and diffuse; (c) grade: low-contrast, slightly desaturated, warmth preserved. Three statements hold; one gets overridden.

Canonical block (tune every bracket to the scene):

Capture Realism: [Foreground subject] sits inside real depth — [thin/light/heavy] atmosphere suspended in the air between camera, subject, and [the far background element], the background rendered softer, desaturated, and lower-contrast than the foreground so the figure sits within the air rather than pasted on a flat plane. [IF WET: Slight moisture has settled on every surface — damp matte hair, slight moisture on skin holding fully matte with no beading and no wet sheen, [wet ground with muted reflection / damp matte fabric / car paint damp but matte not showroom], moisture that mutes and deepens without a single specular hotspot.] Skin reads true cinematic matte — zero shine on forehead, nose bridge, cheekbones, temples, chin, and collarbones, real peach fuzz catching light at the jaw and hairline, real soft fine even pore texture, light absorbed like true subsurface scattering, warmth preserved and natural, slightly desaturated but never pale or washed-out or cool-shifted, never plastic, never doll-skin, never AI-rendered, and never harsh — no acne, no blemishes, no enlarged or rough pores, fine flattering texture that keeps the face looking good. Low-contrast curve — shadows lifted gently holding texture, highlights rolled off softly never clipping to white, nothing crushed to black. All specular highlights surgically removed from skin, hair, fabric, and surrounding surfaces, every pixel reading matte and diffuse. Slightly desaturated grade with warmth preserved.

Tuning: dry → IF-WET-Satz löschen · no humans (M5) → Haut-Satz streichen, Matt-Logik auf Oberflächen · M2 glossy-editorial gewollt → reduzieren/skippen (einziger Mode mit bewusstem Specular) · Atmosphärendichte skaliert (thin Interior, light Exterior, heavy Pre-Dawn/Destroyed) · Block nennt NIE Gear/Hex/fps/Runtime — das ist Camera Capture. Lean positive; die Specular-Kill-/Anti-Plastik-Formeln sind die sanktionierte Negativ-Ausnahme.

CAMERA CAPTURE

Single closing line: body, lens, filter, movement, stock, grade, frame rate, runtime — one trimmed paragraph. The ONLY camera/grade/stock language anywhere in the prompt. Default camera energy is handheld with breath, drift, and organic operator movement — locked-off tripod is OPT-IN only (explicit request or shot type that requires it).

MODE-SELECT TABLE

Mode Use when scene is… Lens Movement Grade
M1 — Narrative Real-world dramatic, lived-in Vintage 2x anamorphic, 40/55/75/100mm, oval bokeh, soft edge falloff Handheld with operator breath Color-negative daylight, fine 35mm grain, teal-amber
M2 — Studio/Editorial White void, clean studio, fashion film, editorial Clean spherical, 32/50/75/100mm, round bokeh, even sharpness Locked tripod + optional slow push Saturated editorial, warm-retained blacks, fine grain; intentional bloom on chrome/rhinestone
M3 — Action/Combat Combat, chase, stunts, debris, smoke Vintage 2x anamorphic, 40/55/75/100mm Handheld and shaky throughout, no stabilized shots Color-negative, heavier low-light grain, palette per scene, dusty haze
M4 — Performance/Concert Stadium, stage, pit, lightstick crowd Vintage 2x anamorphic, streak flares on stage lights Mixed handheld pit-photographer + orbital, hard cuts Color-negative, fine grain, desaturated cool with warm bloom, stage color cast
M5 — Atmospheric/Empty No-humans plates, landscapes, weather, establishing Vintage 2x anamorphic, 35→85mm push range Locked-off or extremely slow push Color-negative, fine grain, palette-driven (hex per scene)

Camera Capture lines per mode:

M1: Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, soft frame-edge falloff — light diffusion bloom softening highlights, handheld with natural operator breath, color-negative daylight film rendition with fine 35mm grain, teal-amber grade, shallow depth of field, 24fps 180° shutter, [XX] seconds.

M1 multi-shot: Camera Capture: Shot 1 — wide-latitude cinema capture, vintage 40mm 2x anamorphic character at a wide aperture, light diffusion bloom softening highlights, handheld with natural operator breath. Shot 2 — same capture register, 75mm anamorphic character at a wide aperture, low-angle handheld at hip height, tight operator breath. Color-negative daylight film rendition with fine 35mm grain, teal-amber grade, shallow depth of field, 24fps 180° shutter, [XX] seconds total.

M2: Camera Capture: wide-latitude cinema capture, clean spherical [XX]mm character at a wide aperture — natural round bokeh, even sharpness — mild diffusion bloom, locked tripod with optional slow push-in, saturated editorial grade, fine grain, warm-retained blacks, 24fps 180° shutter, [XX] seconds. (+ bei Chrom/Strass: intentional highlight bloom on reflective surfaces, blooming the speculars on chrome and rhinestone.)

M3: Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, soft edge falloff — light diffusion bloom softening highlights, handheld and shaky throughout with no stabilized shots, color-negative film rendition with heavier low-light grain, [palette descriptor] with dusty atmospheric haze, 24fps 180° shutter, [XX] seconds. (Slow-Motion: intercut 96fps high-speed slow-motion at the [moment] holding 180° shutter for natural motion blur.)

M4: Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, horizontal streak flares on stage lights — light diffusion bloom softening highlights, mixed handheld pit-photographer and orbital operator energy with hard cuts between angles, color-negative film rendition with fine grain, [stage-lighting color cast], heavy volumetric haze, real sweat sheen, 24fps 180° shutter, [XX] seconds.

M5: Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, soft edge falloff — light diffusion bloom softening highlights, locked-off or extremely slow push-in only, color-negative film rendition with fine grain, palette grade [hex values], atmospheric haze, weathered material detail, 24fps 180° shutter, [XX] seconds. No humans, environment is the subject.

STACKING MODES (Multi-World Sequences)

Cuts between two worlds (M2-Void ↔ M1-Küche, M3 ↔ M4): write each shot's Camera Capture specs inline in the closing line — don't blend modes into one averaged grade; the cut IS the visual punch. Same-mode sequences: one continuous prompt, hard-cut triggers in Movement, one Camera Capture line with per-shot lens differences.

LENS LENGTH GUIDE

32/35/40mm wide establishing, full-body, group · 50/55mm medium portrait, two-shot, waist-up · 75mm tight portrait, isolation, performance close-up · 85/100mm extreme close-up (eyes, lips, jewelry, fabric). Default: 55mm (M1/M3/M4), 50mm (M2); M5 eher 35→55mm.

FRAME RATE NOTES

All modes default 24fps with 180° shutter. Slow-motion beats: intercut 96fps high-speed slow-motion at [moment] holding 180° shutter. — base stays 24fps.

RUNTIME & PER-SHOT TIMING

Runtime stated in three places (title, Frame Map bei Sequenzen, Camera Capture) — all must match. Always ask, never default. Guidance: 48s one strong action · 812s action + reveal/hold · 1215s 23 beats with hard cuts · complex sequences → separate prompts. Per-shot timing must sum to total.

NEGATIVE → POSITIVE REWRITES

Instinct (negative) Lock (positive)
Don't change face @image1 keeps the same face, hair, wardrobe, and silhouette throughout.
Don't switch positions @image1 remains in the left third throughout; @image2 remains in the right third throughout. Neither crosses the center line.
Don't drift Boots stay planted on the same ground marks across the full runtime. Only breath, eyes, hair, and fabric move subtly.
Don't change costume Wardrobe identical across the runtime.
No extra people The frame contains only @image1 and @image2 in their specified positions. No other figures enter or pass through.
No on-screen text No on-screen text, no captions, no signage typography, no rendered text in the frame.
No camera chaos Slow controlled handheld with natural operator breath, preserving @image1 in the left third and @image2 in the right third throughout.
No blur Subjects remain sharply focused; controlled cinematic motion blur appears only on falling rain and distant background light sources.
Don't blink mid-action Gaze stays locked on @image2 across the full runtime, eyes steady, no break in eye contact.
No mode switching The shot runs as one continuous take with no cuts, no scene change, no time jump.

Always prefer the positive form. Negative phrasing only in the explicit suppression lines for known failure modes.

PRE-DELIVERY PASS (Silent QA)

Character gate asked and carried · every reference listed in check + bullet list + @imageN, order matching · canonical reference for every named subject (never substituted by the plate) · mode selected with rationale · Frame Map written · Subject Lock per character (wardrobe NOT re-described) · Cross-Frame Rules if 2+ · Movement with four layers and timestamps · Last Frame with suppression line · World Plate · Sound Bed diegetic · Capture Realism tuned (wet only if wet, skin only if humans, no gear overlap) · Camera Capture single line at bottom · lens chosen · runtime confirmed and matching · per-shot timing sums · no names/brands/tool names/context/meta · no music · English only · three-part delivery · all ten blocks in exact order · every reference tagged · negatives rewritten positive · word count in range.

Repair pass: too poetic → physical visual instructions · overloaded → split into sequence · drift risk → tighten Subject Lock (contact points, ground marks) · swap risk → tighten Cross-Frame Rules · wardrobe re-described → cut, trust the reference · double camera spec → collapse to one line · mode conflict → one dominant mode · action too complex → one dominant motion · Last Frame vague → specific closing composition · over word count → trim Subject Lock and Movement first, then Cross-Frame Rules.

OPTIONAL HANDOFF — BANANA PRO DIRECTOR

If the user pairs a Seedance prompt with an existing Banana Pro plate: ask which cinema mode the plate used and lock the matching camera grammar — the two skills share the same five-mode framework; paired, still and video share visual DNA. Otherwise don't bring it up.