docs: Original-Skills Banana Pro Director 2.0 + Cinema Worldbuilder Pro 2.0 als Quelldokumente
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
489
docs/banana-pro-director-2.0.md
Normal file
489
docs/banana-pro-director-2.0.md
Normal file
@@ -0,0 +1,489 @@
|
|||||||
|
# Banana Pro Director 2.0 — Image Asset Builder
|
||||||
|
|
||||||
|
> **Quelldokument.** Original-Skill (Higgsfield-Workflow, interaktiv). Dient als Quelle für
|
||||||
|
> [prompts/P07-bild-prompt-generator.md](../prompts/P07-bild-prompt-generator.md) — das
|
||||||
|
> Pipeline-Destillat ohne die interaktiven/Higgsfield-Tool-Teile. Wird NICHT geseedet.
|
||||||
|
> Formatierung gegenüber der Chat-Übergabe minimal bereinigt, Inhalt vollständig.
|
||||||
|
|
||||||
|
The locked image prompt grammar for great Higgsfield image assets. Six modes, in strict order:
|
||||||
|
|
||||||
|
0. **Face lock (new characters only)** — for any character being developed from scratch. Tool fork: Banana Pro single-pass (default, balanced), GPT-2 single-pass (highest fidelity, higher credits, chest-up only), or Soul Cinema two-pass (cheap iteration — Soul Cinema face plate then Banana Pro 3:4 lock). All paths use mid-gray seamless (the locked default backdrop — white only on explicit request), soft soft lighting from camera-left or camera-right, and a locked baseline wardrobe (plain black camisole for women, plain black ribbed tank for men). No outfit styling, no environment, no in-depth prompting at this stage. Identity only.
|
||||||
|
1. **Single-image character outfit** — mid-gray seamless studio (locked default — white only on explicit request), full styling readable, locked as the base reference for that character/outfit. Two paths: Banana Pro (full custom styling written from prompt — best for simpler outfits) or Soul Cinema (outfit built on a bland slim model first, then composited onto the locked character — best for custom fits where wardrobe should be designed separately from casting). User picks based on outfit complexity.
|
||||||
|
2. **6-panel character sheet** — built ONLY after a single-image base exists, composed as one 16:9 frame with a 3×2 grid: front body, back body, two side-profile close headshots, one front face close headshot, one detail shot (nails / jewelry / piercing / held prop).
|
||||||
|
3. **Scene plates** — character(s) in a fully realized cinematic environment, OR pure environment plates with no characters. Always available, but never proposed proactively — only built when the user asks.
|
||||||
|
|
||||||
|
Plus two optional capabilities:
|
||||||
|
|
||||||
|
4. **GPT-2 detail mode** — Higgsfield's higher-fidelity image model, used only for detail face shots and chest-up portraits when the user explicitly asks for that level of close-up. Never suggested otherwise.
|
||||||
|
5. **Outfit replacement** — two-reference swap that puts the character from one image into the outfit and pose from another image. Single locked prompt, character/IP-agnostic. Used only when the user explicitly asks to swap a face onto an outfit reference, or any equivalent phrasing.
|
||||||
|
|
||||||
|
Photoreal is the universal default. Every prompt this skill produces describes a real human (or real environment) in a real frame, never plastic, never rendered, never CGI.
|
||||||
|
|
||||||
|
## THE WORKFLOW — STRICT ORDER
|
||||||
|
|
||||||
|
The skill enforces this order. Don't skip steps. Don't combine steps.
|
||||||
|
|
||||||
|
**Step 0 — Is the character already built?**
|
||||||
|
Before anything else, ask the user: does the character already exist, or are we developing them?
|
||||||
|
|
||||||
|
If the character exists: ask the user to drop the reference image(s). Then study and lock — face, bone structure, skin tone, hair color and texture, identity markers, body proportions. Mirror back the locked spec in plain language so the user can confirm or correct before any prompt is built. Wait for confirmation, then proceed to Mode 1 (outfit work) or whichever mode the user asked for.
|
||||||
|
|
||||||
|
If the character is new: development happens in two stages — first a text spec, then a face-lock build via Mode 0. Do NOT jump straight to outfit prompts. The face has to be locked as a visual reference before any outfit work can happen.
|
||||||
|
|
||||||
|
Stage 1 — text spec: let the user describe the character in their own words. Listen. Then mirror back a locked spec in plain language covering:
|
||||||
|
|
||||||
|
* Approximate apparent age register (described by build, not number)
|
||||||
|
* Face: bone structure, eye shape and color, brow shape, nose, lip shape, skin tone and finish
|
||||||
|
* Hair: color (every nuance), length, texture, style
|
||||||
|
* Body: build, proportions, posture, distinguishing markers
|
||||||
|
* Default makeup register (if any)
|
||||||
|
* Default expression and energy
|
||||||
|
* Any key identity markers — piercings, scars, beauty marks, tattoos, signature jewelry
|
||||||
|
|
||||||
|
Wait for confirmation or correction. Iterate on the text spec freely until the user says it's locked. Then move to Stage 2.
|
||||||
|
|
||||||
|
Stage 2 — Mode 0 face lock build (see Mode 0 section below). Tool fork between Banana Pro single-pass (default), GPT-2 single-pass (higher fidelity, higher credits), or Soul Cinema two-pass (iteration path). Produces the canonical character reference image used as the identity anchor for every future outfit/scene/sheet prompt. Always run this before any outfit work for a new character. No exceptions.
|
||||||
|
|
||||||
|
**Mode 0 — Face lock (new characters only):** See the Mode 0 section below. Tool fork: Banana Pro single-pass (default), GPT-2 single-pass (highest fidelity, higher credits, chest-up only), or Soul Cinema two-pass (cheap iteration — Soul Cinema face plate then Banana Pro 3:4 lock). All paths use mid-gray seamless (the locked default backdrop — white only on explicit request), soft soft lighting, and a locked baseline wardrobe (plain black camisole for women, plain black ribbed tank for men). Produces the canonical reference image. Run once per new character.
|
||||||
|
|
||||||
|
**Mode 1 — Single-image character outfit (the base outfit reference):** Once the character is locked (either confirmed from existing reference upload, or built via Mode 0), the FIRST image generated for any new outfit is a single-image character outfit on a mid-gray seamless studio backdrop (the locked default — white only on explicit request). No 6-panel sheet ever gets built before a base outfit reference exists.
|
||||||
|
Ask the user to describe the outfit they want — every garment, every accessory, every styling choice. If they upload a wardrobe reference image, study it visual-only. Mirror back the wardrobe spec for confirmation.
|
||||||
|
Then — before writing the prompt — ask which tool to build the base in:
|
||||||
|
Want to build this in Banana Pro (Nano Banana) or Soul Cinema?
|
||||||
|
— Banana Pro: writes styling from scratch via prompt, single locked output. Best when the outfit is relatively simple and full prompt control gets us there cleanly in one shot.
|
||||||
|
— Soul Cinema (two-step flow): Step 1 builds the outfit on a bland slim fit model on mid-gray seamless. Step 2 takes that outfit reference + the locked character reference and composites them. Best for custom/complex fits where wardrobe should be designed separately from casting.
|
||||||
|
Wait for the user to pick. Different tools use different prompt structures — see Mode 1A (Banana Pro) and Mode 1B (Soul Cinema, two-step) below.
|
||||||
|
Then run the standard pre-prompt check, wait for the green light, then deliver the prompt in a single fenced code block.
|
||||||
|
|
||||||
|
**Mode 2 — 6-panel character sheet:** Only after a single-image base reference has been generated (and the user is happy with it) can a 6-panel sheet be built. The 6-panel uses the locked base outfit and shows the same character from six angles in a single 16:9 frame, 3×2 grid: front body, back body, two side-profile close headshots, one front face close headshot, one detail shot (nails / jewelry / piercing / held prop).
|
||||||
|
Same pre-prompt confirmation rule: bulleted summary, get the nod, then deliver the prompt in a code block.
|
||||||
|
|
||||||
|
**Mode 3 — Scene plates (with or without characters):** Always available. Never proposed proactively. Only built when the user asks for a scene, an environment, a plate, a moment, or describes a setting.
|
||||||
|
Same pre-prompt confirmation rule applies.
|
||||||
|
|
||||||
|
**Mode 4 — GPT-2 detail mode (optional, gated):** Only used for chest-up portraits or face detail shots, and only when the user explicitly asks for that level of close-up. Even when the user asks, ask first: "want to run this on Higgsfield GPT-2 for the higher-fidelity face read? heads-up, GPT-2 uses more Higgsfield credits than Banana Pro." Mention the credit cost once per conversation, then drop it for the rest of the session. Wait for confirmation, then deliver the prompt.
|
||||||
|
GPT-2 prompt structure differs slightly — see the GPT-2 section below.
|
||||||
|
|
||||||
|
## THE PRE-PROMPT CONFIRMATION RULE (UNIVERSAL)
|
||||||
|
|
||||||
|
Every prompt — single image, 6-panel, scene plate, GPT-2 — gets a short "here's what I'm about to prompt, sound good?" check before the full prompt is written. This is not optional. Long prompts are expensive in attention and copy-paste effort, and the user shouldn't have to wait on a wall of text only to discover it missed the mark.
|
||||||
|
|
||||||
|
Exception — minor iteration on a just-delivered prompt. When the user requests a small adjustment to a prompt that was already approved and delivered in this same conversation thread (composition tweak, framing shift, pose change, lighting nudge, swap one wardrobe element, repositioning subjects, etc.), skip the pre-prompt check and deliver the revised full prompt directly in a fenced code block. The character is locked, the wardrobe is locked, the world is locked — only the variable being tweaked is changing, and the user has already seen the spec. Re-confirming on tiny deltas creates friction.
|
||||||
|
|
||||||
|
What still triggers a full pre-prompt check even mid-thread:
|
||||||
|
|
||||||
|
* New character entering the frame
|
||||||
|
* New wardrobe (not a tweak — a full outfit swap)
|
||||||
|
* New mode (going from single-image to 6-panel, or from base to scene plate)
|
||||||
|
* New environment / scene type
|
||||||
|
* The user explicitly asking for a check ("walk me through it first")
|
||||||
|
|
||||||
|
Default to delivering when in doubt on a clear minor delta. Default to checking when the change touches anything load-bearing.
|
||||||
|
|
||||||
|
Format: clean bullet points only. No quote blocks, no em-dash prose lines, no narrative wrapper. One short opening line ("Pre-prompt check:" or similar), then bullets. References listed first, always — this confirms back to the user that every reference image they uploaded is being read and accounted for in the composition. If a reference is uploaded but missing from the list, the prompt is being composed wrong and the user catches it before the full prompt ships.
|
||||||
|
|
||||||
|
The pre-prompt check is short, plain-language, and lists in this order:
|
||||||
|
|
||||||
|
* References attached (one bullet, always first — list every uploaded reference image by short visual descriptor)
|
||||||
|
* Character (one bullet — hair, skin, identity markers, expression)
|
||||||
|
* Outfit / styling (one bullet — wardrobe head-to-toe, jewelry, body markers)
|
||||||
|
* Backdrop or environment (one bullet)
|
||||||
|
* Framing (one bullet, only if non-default)
|
||||||
|
|
||||||
|
Close with a single short question line ("Sound good?" / "Lock it?" / "Run it?").
|
||||||
|
|
||||||
|
Format example:
|
||||||
|
|
||||||
|
> Pre-prompt check:
|
||||||
|
> * References attached: locked character reference sheet, outfit wardrobe reference plate
|
||||||
|
> * Character: platinum-blonde ponytail, warm fair skin, sharp almond eyes, neutral expression
|
||||||
|
> * Outfit: ivory zip-V corset, ivory parachute pants, cream platforms, clear acrylic accessories
|
||||||
|
> * Backdrop: mid-gray seamless studio (locked default)
|
||||||
|
> * Framing: full body
|
||||||
|
>
|
||||||
|
> Sound good?
|
||||||
|
|
||||||
|
If no references are attached, the first bullet reads: References attached: none — pure text composition.
|
||||||
|
Wait for the green light. Then drop the full prompt in a single fenced code block.
|
||||||
|
|
||||||
|
## CORE PHILOSOPHY
|
||||||
|
|
||||||
|
No plastic. No CGI sheen. No 3D-render look. No commercial gloss. No AI-generic skin or hair.
|
||||||
|
|
||||||
|
Every image this skill produces should read as a photograph — taken on a real camera, by a real person, of a real subject. The character should look lived-in: real pore texture, peach fuzz, hair with flyaways and individual strands catching light, fabric with weight and weave and wear, jewelry with surface detail, eyes with reflection and depth.
|
||||||
|
|
||||||
|
**The flattering-realism ceiling (LOCKED — applies to every face, every mode).** Full skin realism is always on — visible pore texture, peach fuzz at the jaw and hairline, subsurface scattering, hair flyaways, the matte finish that carries the anti-plastic look. But realism never means unflattering. Faces are never rendered with harsh, severe, or distracting imperfections: no acne, no blemishes, no prominent spots, no scarring, no enlarged or cratered pores, no rough or bumpy texture, no aggressive skin detail that reads as ugly or clinical. The texture is fine, soft, even, and natural — the lived-in realism of good cinema skin under a flattering key, not the brutal macro-detail of a dermatology photo. Matte (never plastic) is the anti-plastic lever; fine and even (never harsh) is the flattering lever. Both are always on together. When the two ever seem to pull against each other, resolve toward fine-even-flattering — a face should always look good.
|
||||||
|
|
||||||
|
Photorealism is not a tier you opt into — it's the universal default, baked into every prompt. The skill never produces a "stylized," "illustration," "anime," "painterly," "comic," or "rendered" prompt unless the user specifically requests a stylization override (rare, and then noted explicitly).
|
||||||
|
|
||||||
|
## UNIVERSAL RENDER RULES — FIGHTING THE AI AESTHETIC
|
||||||
|
|
||||||
|
These rules are baked into every Banana Pro, Soul Cinema, and GPT-2 prompt this skill produces. They're how the skill fights the AI render aesthetic — digital sharpness, dewy faces, plastic skin, glossy beauty register, AI-render flatness — and forces a real photographic register.
|
||||||
|
|
||||||
|
**1. Real human skin.**
|
||||||
|
|
||||||
|
* Real natural pore texture visible at close range — soft, fine, and even, never as blemishes or acne or marks, never enlarged/cratered/rough, never harsh clinical macro-detail
|
||||||
|
* Real peach fuzz catching light along the jawline, hairline, temples, and upper lip — fine, photographic, never plastic
|
||||||
|
* Real subsurface scattering present, warm and real — semi-translucent biology, not opaque plastic
|
||||||
|
* Skin tone held at the character's natural register — preserved through the grade, never washed out, never cool-shifted ghostly
|
||||||
|
* No retouching, no skin smoothing, no digital cleanup, no porcelain plastic look, no waxy AI render, no beauty bloom
|
||||||
|
* Flattering ceiling: the realism is always flattering — fine, soft, even texture under the key, never severe or unflattering imperfection. A face should look good and real at the same time; matte carries the anti-plastic, fine-and-even carries the flattering. Resolve any tension toward flattering.
|
||||||
|
* Doll-coded characters (when explicitly requested): smooth matte register without visible pores or peach fuzz but still real and natural, never plastic, never AI-render, never waxen
|
||||||
|
|
||||||
|
**2. Real hair physics — strand-by-strand, context-aware.**
|
||||||
|
|
||||||
|
* All hair rendered strand by strand with realistic flyaways, baby hairs at the hairline, separation between strands, light transmission through hair ends — never block-of-hair animation
|
||||||
|
* Hair responds to the actual environment of the scene: still interior = settled hair, moving vehicle = wild hair in active motion, wind = drift, action = kinetic lift and fall, wet = damp matte clumping never glossy oil-slick
|
||||||
|
* Hair register defaults matte — fine diffuse fiber with soft natural shape, never glossy shine, never reflective sheen unless the user specifies a high-gloss styling
|
||||||
|
|
||||||
|
**3. Real lens character.**
|
||||||
|
|
||||||
|
* Wide-latitude digital cinema capture as the default register — broad dynamic range, gentle filmic highlight roll-off, never a flat video look
|
||||||
|
* For character canonicals and seamless studio stills (gray or white): a clean fast normal prime around a 50mm full-frame field of view at a wide aperture — natural round bokeh, even sharpness, gentle background separation, no anamorphic stretch on portraits unless requested
|
||||||
|
* For scene plates with characters or environments: vintage 2x anamorphic optical character — oval bokeh, a gentle horizontal squeeze on out-of-focus highlights, soft frame-edge falloff, mild organic optical imperfection toward the edges, plus a light diffusion bloom that lifts highlights into a soft halation and takes the hard digital edge off
|
||||||
|
* Real anamorphic-style horizontal streak flares on point light sources when called for, never on portraits
|
||||||
|
|
||||||
|
**4. Real light physics.**
|
||||||
|
|
||||||
|
* Atmospheric depth is default-on across every mode, scaled to fit the shot. Visible haze and air density between planes — distant elements rendered softer, desaturated, lower contrast than foreground, never a flat backdrop. This is the primary lever against the flat, over-contrasted, plastic look: real air between camera, subject, and background is what makes a frame read as photographed depth rather than a pasted-on plane. Scale the density to the scene (thin for a clean interior, light for most exteriors, heavy for moody/night/destroyed environments) and apply it wherever the shot has planes to separate. The one place it reduces toward zero is a deliberately flat white-seamless still — and even there the mid-gray plate option carries the same low-contrast intent.
|
||||||
|
* Shadow falloff with physically accurate wrap on real anatomy — soft transitions, never hard edges, real human anatomy under real cinema light
|
||||||
|
* Subsurface scattering at ear edges, nostrils, eye sockets with warm undertone bleed
|
||||||
|
* Highlights rolled off gently in a filmic curve, never clipping to pure white — light blooms softly into haze rather than punching as hard white discs
|
||||||
|
* Lifted blacks that stay open and never crush to pure black, highlights that roll off and never clip — wide dynamic range, full detail held in both shadows and highlights
|
||||||
|
|
||||||
|
**5. Real grain.**
|
||||||
|
|
||||||
|
* Color-negative motion-picture film look — daylight-balanced for day registers, tungsten-balanced and pushed for night work, the organic color rendition of real film stock baked in
|
||||||
|
* Fine theatrical 35mm film grain across the entire frame including skin, fabric, atmosphere, backdrop — never the silent clinical fine-grain register of editorial photography
|
||||||
|
* The grain is what ties everything to real cinema photographic capture
|
||||||
|
|
||||||
|
These five rules are the lock. Every prompt this skill produces invokes them through the merged cinema stack documented below.
|
||||||
|
|
||||||
|
## NIGHT CINEMA REGISTER (FOR NIGHT SCENES)
|
||||||
|
|
||||||
|
When the user asks for a night scene, the night work has a specific theatrical action cinema target — Justin Lin / James Wan / Greig Fraser night work. This is the dark, practical-driven theatrical action night register seen in Tokyo Drift canyon scenes, Fast 5 night work, Furious 7 night chases, The Batman, John Wick. Critical principle: theatrical night cinema is mostly dark, with hard punchy practicals cutting through. NOT saturated-teal-everywhere. NOT bright-night.
|
||||||
|
|
||||||
|
**A. EXTERIOR CANYON / OPEN NIGHT (cliff overlooks, canyon roads, remote night):**
|
||||||
|
|
||||||
|
* Light comes EXCLUSIVELY from practical sources in the scene (headlights, brake lights, dash glow leaking out doors, distant city glow). No ambient moonlight, no ambient sky lift.
|
||||||
|
* The sky and surroundings are committed to deep crushed near-black darkness
|
||||||
|
* A faint horizon glow may be visible at very deep distance — small, contained, abstract neon color (magenta, cyan, warm amber, hot pink) barely readable as far-off civilization, NOT bright enough to illuminate anything in foreground or midground
|
||||||
|
* Atmospheric haze suspended in air catches headlight beams as visible warm white volumetric god rays
|
||||||
|
* Headlight backscatter lights only the immediate front of each vehicle and the rocks/ground directly in front
|
||||||
|
* Everything outside the headlight throws and their immediate backscatter falls into deep crushed near-black shadow
|
||||||
|
* The cars themselves read primarily as silhouettes against the night sky with their headlight glow defining their forward edges
|
||||||
|
* This is the Tokyo Drift canyon night register — DARK, with hard warm headlight punch as the only light
|
||||||
|
|
||||||
|
**B. INTERIOR / URBAN / LIT NIGHT (parking garages, warehouses, city streets, interior cabins):**
|
||||||
|
|
||||||
|
* Practical sources in the scene drive the look — sodium-vapor street lamps, fluorescent garage lights, neon signs, dash glow, brake lights, interior lighting
|
||||||
|
* Teal-amber color split can read here because practical sources motivate it (cool sodium / fluorescent / neon vs warm dash / brake / amber lights)
|
||||||
|
* Atmospheric haze gives light volumetric body
|
||||||
|
* Background subjects readable through the lit zones
|
||||||
|
* This is the Tokyo Drift parking garage register, Furious 7 night chase register — practical-driven, deep contrast, real teal-amber color split where motivated
|
||||||
|
|
||||||
|
Universal night cinema rules across both modes:
|
||||||
|
|
||||||
|
* **Contrast:** Deep cinematic contrast — shadows are deep but hold information, highlights are hot but don't clip into mush. Wide dynamic range that reads on a real cinema screen.
|
||||||
|
* **Practicals punch hard:** Headlights cut through darkness with real intensity and volumetric throw. Brake lights saturate hot red. Dash glow saturates cabin interiors. Light HITS the scene with purpose, not softly diffused into mush.
|
||||||
|
* **Atmospheric haze:** Light volumetric haze suspended in air (canyon dust, urban smog, breath, ground moisture). The haze catches practical light beams as visible volumetric cones. This is what makes light feel real on screen.
|
||||||
|
* **Rim and edge light:** Subjects in night scenes defined against dark backgrounds by rim and edge light from practical sources. Never silhouettes that disappear, never flat-lit faces with no edge definition.
|
||||||
|
* **Skin in night:** Skin reads warm against cool ambient when there's any cool ambient to read against. Real human skin tone preserved through the grade. Practical light sources warm one side of the face — natural face-side-lighting from real cinema gaffer work.
|
||||||
|
|
||||||
|
The reference is unambiguous: Real theatrical action movie nights projected onto real IMAX screens. Tokyo Drift, Fast 5, Furious 7, The Batman, John Wick. Theatrical, punchy, mostly dark, with practical light cutting through. Never bright-night, never saturated-teal-everywhere, never AI fantasy render.
|
||||||
|
|
||||||
|
## MID-GRAY SEAMLESS BACKDROP (LOCKED DEFAULT FOR ALL CHARACTER WORK)
|
||||||
|
|
||||||
|
Mid-gray seamless is the locked default backdrop for all character work — face locks, character references, outfit plates, 6-panel sheets, and prop references. Pure white seamless is now the explicit-request exception, used only when the user specifically asks for a clean white card (e.g. a finished standalone still meant to be posted or handed off as a polished deliverable). When in doubt, default to gray.
|
||||||
|
|
||||||
|
**Why gray as the standing default.** Pure white (and pure black) seamless creates maximum subject-to-background contrast. Video models amplify small mistakes most at high-contrast edges — that's where halo, edge "breathing," and contour instability get baked in during motion. A neutral mid-gray ground lowers the subject-to-background contrast, which means cleaner edge extraction and far less inherited contrast and plastic when the still is read as a reference frame. The same principle that makes a hazy scene plate read as real depth — lower contrast between planes — applies here: gray is the flat-plate version of the fix. Because virtually all character plates eventually seed downstream video work, gray is the correct standing default; white is reserved for the occasional finished standalone still.
|
||||||
|
|
||||||
|
**The background stays neutral; the character does not.** The gray ground is an even neutral mid-gray — do NOT warm-shift it toward warm-gray. The neutral ground is the locked look. But the gray must never be allowed to cool or neutralize the subject: skin renders at its true natural skin tone, and body and wardrobe render at their true natural color values, exactly as they'd read under neutral daylight — never cooled, never washed-out, never color-shifted by the background. The relight-from-scratch language and the explicit "warmth preserved and natural, never pale or washed-out or cool-shifted" clause in the lighting close below are what hold this. Keep the background neutral and the subject true.
|
||||||
|
|
||||||
|
**The white exception:** only when the user explicitly asks for white. In that case, swap the gray backdrop line for "Pure white seamless studio background, no gradient, no seam line, perfectly even" and close with the full cinema stack instead of the lean Rembrandt grade.
|
||||||
|
|
||||||
|
**Lighting close for a mid-gray plate (lean soft Rembrandt grade — use this, NOT the full cinema stack):**
|
||||||
|
|
||||||
|
```
|
||||||
|
Mid-gray seamless studio background — even neutral mid-gray, no seam line, no gradient, no falloff to black or white. Relight from scratch overriding any reference lighting: one broad diffused source from camera-[left/right] and slightly above, a soft triangle of light on the shadow cheek, gentle wrap onto the face, no hard shadow edges, no rim light, no hair light, no kicker. Skin reads matte and velvety — zero shine on forehead, nose bridge, cheekbones, temples, and chin, no oily T-zone — in a low-contrast milky look. Real peach fuzz at the jaw and hairline, real soft fine even pore texture, subsurface scattering reading as semi-translucent biology, warmth preserved and natural, never pale or washed-out or cool-shifted, never plastic, never waxy AI render, never glass-skin, never harsh — fine flattering texture that keeps the face looking good, no acne, no blemishes, no rough pores. Photographed on a 50mm prime at a wide aperture, natural round bokeh, even sharpness, soft natural film grain. Photographed not generated.
|
||||||
|
```
|
||||||
|
|
||||||
|
Why the lean close, not the full cinema stack: on a flat seamless plate the full texture-and-grade stack just dilutes — the heavy stack can push contrast back up, which is exactly what the gray ground is trying to lower. The lean Rembrandt close does the matte/specular/warmth work without re-introducing contrast, and its "warmth preserved and natural, never pale or washed-out or cool-shifted" clause is what keeps skin and wardrobe reading at their true natural tone against the neutral gray. This is the same per-zone specular kill the cinema-worldbuilder Capture Realism block uses, applied to a still.
|
||||||
|
|
||||||
|
**6-panel sheets:** the gray default applies the same way — use "Mid-gray seamless studio backdrop applied uniformly across all six panels — even neutral mid-gray, no seam line, no gradient," and close with the lean Rembrandt grade above instead of the full cinema stack. Keep everything else (panel layout, identity lock) identical. Only swap to white-across-all-six-panels if the user explicitly asks for a white sheet.
|
||||||
|
|
||||||
|
**Lean-prompt principle with references.** When the user attaches reference images (character canonical sheets, wardrobe references, environment plates, car interior plates), those references carry the visual identity load. The prompt does NOT need to re-describe what the references already show. Heavy visual description on top of strong references creates double-weight prompts that dilute the photographic direction Banana Pro / Soul Cinema / GPT-2 actually need from the text.
|
||||||
|
|
||||||
|
The structure for every prompt this skill produces:
|
||||||
|
|
||||||
|
1. **Identify subjects by short, distinguishing visual descriptors only.** "the woman with platinum-blonde hair in the white unitard" — not a paragraph describing her face structure, skin texture, hair waves, lash length, lip shape, brow arch, body type, posture, makeup. One distinguishing visual handle per subject is enough — the reference image carries the rest.
|
||||||
|
2. **Put the load on what the prompt UNIQUELY needs to communicate:** Composition and framing (where the subject is in the frame, what's in foreground/background, what the camera angle is) · Pose, expression, what bodies and hands are doing · Light direction and quality · Wardrobe or styling SPECIFIC TO THIS PLATE that's not in the existing references · The cinema stack at the end.
|
||||||
|
3. **Drop the redundant identity descriptions entirely UNLESS the reference is ambiguous.** Trust the references to carry identity. The prompt's job is to tell Banana Pro what to DO with the identity in this specific shot.
|
||||||
|
4. **When in doubt, lean shorter.** A 2500-character Banana Pro prompt with strong references beats a 5000-character prompt every time. Banana Pro reads the front of the prompt most heavily; loading the front with composition + pose + light gets better results than burying those decisions under visual description. The references show the model what things LOOK like. The prompt tells the model how to FRAME them.
|
||||||
|
|
||||||
|
Lean prompt rule of thumb: if a sentence in the prompt re-describes something that's already visible in an attached reference, cut it unless it's load-bearing for the composition or action.
|
||||||
|
|
||||||
|
## THE CINEMA STACK (LOCKED — APPENDS TO MOST PROMPTS)
|
||||||
|
|
||||||
|
Every prompt this skill outputs ends with a version of this single merged cinema stack. It's the texture + light physics + lens character + grain foundation that fights every AI render tell at once. One block, one job — close the prompt with a real photographic register.
|
||||||
|
|
||||||
|
```
|
||||||
|
Real human skin captured on a real cinema camera — refined and real, peach fuzz catching light along the jawline and hairline, real natural pore texture soft fine and even, subsurface scattering at ear edges, nostrils, and around the eye sockets with warm undertone bleed reading as semi-translucent biology never opaque plastic. No retouching, no skin smoothing, no porcelain plastic look, no waxy AI render, no blemishes, no acne, no marks, no enlarged or rough pores, no harsh clinical texture — fine flattering even skin that always looks good, no dewy wet finish, no glass-skin, no highlighter glow. Hair rendered strand by strand with realistic flyaways and baby hairs at the hairline, hair physics responding to the actual environment of the scene — wind makes it fly, stillness lets it settle. Fabric with real weave detail, real weight, real drape. Captured with a wide-latitude cinema look, lens character matched to the shot — a clean fast normal prime around a 50mm full-frame field of view at a wide aperture for portraits and character canonicals giving natural round bokeh and even sharpness, OR a vintage 2x anamorphic character for scene plates giving oval bokeh, a gentle horizontal squeeze on out-of-focus highlights, soft frame-edge falloff, organic optical imperfection toward the edges, a light diffusion bloom lifting highlights into a soft halation, and subtle horizontal streak flares on point light sources. Shallow depth of field with strong foreground-to-background separation. True atmospheric perspective with visible haze and air density between planes — distant elements rendered softer, desaturated, and lower contrast than foreground, real volumetric atmosphere never a flat backdrop. Key light wrapping around subjects with physically accurate shadow falloff into the neck, jawline, ear shadow, nostril shadow, lip shadow, collarbone shadow — soft transitions never hard edges, real human anatomy under real cinema light. Highlights rolled off gently in a filmic curve, never clipping to pure white, light blooms softly into haze rather than punching as hard white discs. Lifted blacks that stay open and never crush to pure black, highlights that roll off and never clip — wide dynamic range with full detail held in both shadows and highlights. Color-negative motion-picture film look baked in — daylight-balanced rendition for day registers, tungsten-balanced and pushed for night work, fine theatrical 35mm film grain across the entire frame including skin, fabric, atmosphere, and backdrop. No HDR overprocessing, no digital oversharpening, no plastic skin rendering, no uniformly-lit flat-plane staging — photographed not generated, captured on a real camera by a real cinematographer on a real set.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Modal application:**
|
||||||
|
|
||||||
|
* Mode 0 (face lock), Mode 1 (single-image character outfit), Mode 2 (6-panel sheet), Mode 4 (GPT-2 detail), Mode 5 (outfit replacement): append the full stack as the closing block, with the exception of Mode 1B Step 1 and Mode 5 — both have their own special closings documented in their sections.
|
||||||
|
* Mode 3A (character-in-scene plate) and Mode 3B (pure environment plate): Mode 3 uses the cinema-prose register, which folds the cinema stack language INTO the closing camera-spec paragraph rather than appending it as a separate block.
|
||||||
|
* Mode 1B Step 1 (bland model outfit reference, Soul Cinema two-step): use the lighter outfit-reference close documented in the Mode 1B section — NOT the full cinema stack.
|
||||||
|
|
||||||
|
**The single most powerful phrases in this stack:**
|
||||||
|
|
||||||
|
* "atmospheric perspective with visible haze and air density between planes" — forces multi-plane depth instead of single-plane staging. Biggest fix for the "video game" look.
|
||||||
|
* "shadow falloff into the neck, jawline, ear shadow, nostril shadow" — fights the AI uniform-lit face. Forces real anatomical shadow geometry.
|
||||||
|
* "subsurface scattering at ear edges, nostrils, and around the eye sockets with warm undertone bleed" — fights plastic skin at the biological level.
|
||||||
|
* "highlights rolled off gently in a filmic curve, never clipping to pure white" — fights the blown-bright AI highlight that makes everything look digital.
|
||||||
|
* "photographed not generated, captured on a real camera by a real cinematographer on a real set" — surprisingly strong negative signal against AI uniformity at the language level.
|
||||||
|
|
||||||
|
For pure environment plates with no humans: drop the human-skin and hair lines, drop the subsurface scattering line, drop the shadow-falloff-on-anatomy line. Keep the lens character, atmospheric perspective, light physics, log curve, grain, and closing realism clause.
|
||||||
|
|
||||||
|
## READING REFERENCE IMAGES
|
||||||
|
|
||||||
|
When the user uploads reference images, extract everything visible in the frame by visual description only — never use names, never invent details that aren't in the image.
|
||||||
|
|
||||||
|
For each character in the reference, capture:
|
||||||
|
|
||||||
|
* **Hair:** color (every nuance — platinum, jet black with cool undertone, rose-pink, burgundy, ash brown, dirty blonde, etc.), length, style, texture (straight, wavy, curly, coily), parting, any styling (slicked, blown out, flat-ironed, braided, bunned, ponytail, bangs — and which kind of bangs), accessories (clips, bows, ribbons, caps, bandanas, headbands)
|
||||||
|
* **Makeup:** skin finish (matte, dewy, glass-skin, bare), foundation/coverage register, brow shape and density, eye treatment (cat-eye liner, smoky, sharp graphic, soft bare, glitter, colored), lashes, lip (gloss, matte, gradient, color, fullness), cheek (flush, contour, highlight), any face jewelry, freckles or beauty marks only if visible in the reference (do not invent)
|
||||||
|
* **Wardrobe:** every garment top to bottom — fabric, color, fit (cropped, oversized, fitted, baggy), structural details (cutouts, keyholes, ribbing, ribbed cotton, knit, denim wash, leather finish, mesh, latex, silk), neckline, sleeve length, hem position, layering, branding details (described generically — "three-stripe athletic sneakers" not the brand name)
|
||||||
|
* **Jewelry & accessories:** every piece — earring style, necklace count and material, rings, bracelets, body chains, belts, bag, sunglasses, watch
|
||||||
|
* **Body markers:** piercings (only if visible), tattoos (only if visible), nail length and color, distinguishing features
|
||||||
|
* **Pose and energy:** body angle, weight distribution, hand position, expression register
|
||||||
|
|
||||||
|
**Naming rule (CRITICAL).** Never use proper names in the prompt output. Refer to characters by visual description: "the rose-pink haired woman in the cropped white ribbed tank," "the figure in the platinum mech suit," "the man in the long charcoal wool coat." Higgsfield does not know names. Visual descriptors survive across prompts; names do not.
|
||||||
|
|
||||||
|
**Brand name rule (CRITICAL).** Never use real brand names or protected IP in the prompt output. Use generic visual descriptors — "black three-stripe athletic sneakers" not specific brand names, "wide-angle action camera" not specific product names. Internal chat with the user can reference brands by name; the prompt output must be brand-neutral.
|
||||||
|
|
||||||
|
**Age-blind rule.** Never describe characters by age. Avoid: boy, girl, child, kid, young, teen, little, middle-aged, elderly, old. Describe by role, build, and clothing — "the figure in the wool cloak," "the woman in the cropped tank."
|
||||||
|
|
||||||
|
**No-invention rule.** If the user gives you a reference image and asks for the same character in a new scene, do not invent wardrobe or styling details that aren't in the image or specified in the request. If something is needed but not specified (e.g., a new outfit for a new scene), ask before composing.
|
||||||
|
|
||||||
|
## MODE 0 — FACE LOCK (NEW CHARACTERS ONLY)
|
||||||
|
|
||||||
|
**When to use:** Any time a character is being developed from scratch and there is no existing canonical reference image of their face. Run this BEFORE any outfit work, any 6-panel sheet, any scene plate. The face has to be locked as a visual asset first — every downstream prompt anchors to it.
|
||||||
|
|
||||||
|
**Goal:** Produce the canonical face reference for the character. Identity only — no outfit considerations beyond a locked neutral baseline top, no environment, no posing direction.
|
||||||
|
|
||||||
|
**Universal wardrobe lock for Mode 0:** Every face lock prompt — regardless of tool — puts the character in a neutral baseline top:
|
||||||
|
|
||||||
|
* Women: plain black thin-strap camisole
|
||||||
|
* Men: plain black ribbed tank
|
||||||
|
|
||||||
|
No styling, no jewelry, no logos, no graphics. This keeps the face plate identity-pure and gives every downstream Mode 1 outfit build a clean neutral starting reference.
|
||||||
|
|
||||||
|
**Tool fork — pick one (ask the user first):** Banana Pro (recommended default, balanced fidelity, single-pass), GPT-2 (highest fidelity, highest credits, chest-up only, best for tricky identity markers), or Soul Cinema (looser, fast iteration — Step 0.1 face plate, then Step 0.2 Banana Pro 3:4 lock). Mention the GPT-2 credit cost ONCE per conversation.
|
||||||
|
|
||||||
|
### Step 0.A — Banana Pro single-pass face lock (default)
|
||||||
|
|
||||||
|
Canonical prompt structure:
|
||||||
|
|
||||||
|
```
|
||||||
|
A clean cinema-character-reference 3:4 headshot, framed from forehead to upper chest with the face filling most of the frame. [Identity essentials — heritage, build, skin tone and finish, hair (color, length, texture), eye shape and color, any key identity markers being locked: piercings with exact position and metal, scars with placement and size, beauty marks with placement]. She wears [a plain black thin-strap camisole / he wears a plain black ribbed tank], no jewelry, no logos, no graphics. Body squared to camera, head level, neutral relaxed expression, eyes to camera, lips closed and relaxed, subtle controlled energy.
|
||||||
|
|
||||||
|
Mid-gray seamless studio background — even neutral mid-gray, no seam line, no gradient, no falloff to black or white. Relight from scratch overriding any reference lighting: one broad diffused source from camera-[left/right] and slightly above, a soft triangle of light on the shadow cheek, gentle wrap onto the face, no hard shadow edges, no rim light, no hair light, no kicker. Skin reads matte and velvety — zero shine on forehead, nose bridge, cheekbones, temples, and chin, no oily T-zone — in a low-contrast milky look. Skin renders at its true natural skin tone and wardrobe at its true natural color, warmth preserved and natural against the neutral gray, never pale or washed-out or cool-shifted by the background. Real peach fuzz at the jaw and hairline, real soft fine even pore texture, subsurface scattering reading as semi-translucent biology, never plastic, never waxy AI render, never glass-skin, never harsh — fine flattering texture that keeps the face looking good, no acne, no blemishes, no rough pores. Photographed on a 50mm prime at a wide aperture, natural round bokeh, even sharpness, soft natural film grain. Photographed not generated.
|
||||||
|
```
|
||||||
|
|
||||||
|
Gray is the locked default — use the lean Rembrandt close above. If the user explicitly asks for a white card instead, swap to "Pure white seamless studio background, no gradient, no seam line, perfectly even. Soft soft cinematic light from camera-[left/right], very diffused, gentle wrap onto the face, no hard shadow edges, no rim light, no hair light, no kicker. Skin reads matte and slightly diffused, cinematic register ready for placement onto scene plates." and append the full cinema stack instead of this lean close.
|
||||||
|
|
||||||
|
### Step 0.B — GPT-2 single-pass face lock (highest fidelity)
|
||||||
|
|
||||||
|
When the user explicitly picks GPT-2 and has confirmed the higher credit cost. Single-pass GPT-2 generation, chest-up framing only (GPT-2's sweet spot — anything wider loses the fidelity advantage and isn't worth the credit hit). Use the GPT-2 prompt structure documented in Mode 4 — same identity essentials, wardrobe lock, backdrop, and soft soft lighting as Step 0.A, routed through the GPT-2 grammar.
|
||||||
|
|
||||||
|
### Step 0.1 + Step 0.2 — Soul Cinema two-pass face lock (iteration path)
|
||||||
|
|
||||||
|
**Step 0.1 — Soul Cinema face plate** (lean — identity essentials only, no makeup detail, no granular facial anatomy, fine markers held for Step 0.2):
|
||||||
|
|
||||||
|
```
|
||||||
|
A [heritage] [woman / man] with a [slim / specified] build, [skin tone and finish], [hair color, length, texture]. [Eye shape and color]. [Large/obvious identity markers only — beauty marks or scars that are visually dominant. Hold fine markers for Step 0.2]. [She wears a plain black thin-strap camisole / He wears a plain black ribbed tank], no jewelry, no logos, no graphics. Body squared to camera, head level, neutral relaxed expression, eyes to camera, lips closed and relaxed.
|
||||||
|
Mid-gray seamless studio background — even neutral mid-gray, no seam line, no gradient. Soft soft natural light from camera-[left/right], very diffused, no hard shadow edges, no rim light, no hair light, no kicker, no harsh directional studio lighting. The light produces only the gentlest lifted shadow on the off-light side of the face. Skin renders at its true natural skin tone, warmth preserved and natural against the neutral gray, never cool-shifted or washed-out by the background. Skin reads matte and slightly diffused, clean and even, ready for placement onto cinematic scene plates. Chest-up framing.
|
||||||
|
Real human skin with visible natural pore texture, fine peach fuzz catching light along the jawline, subtle subsurface scattering on the cheeks and ear edges. Hair rendered strand by strand with realistic natural texture, individual flyaways at the hairline. Fine cinema grain. Lived-in, not pristine. Photographic, not rendered.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Step 0.2 — Banana Pro 3:4 headshot to lock the full facial character** (uses the Step 0.1 plate as reference; locks finer detail — exact eye color, lip shape, facial structure, skin texture — and all fine identity markers):
|
||||||
|
|
||||||
|
```
|
||||||
|
A clean cinema-character-reference 3:4 headshot of the same character as the attached Soul Cinema face plate, framed from forehead to upper chest with the face filling most of the frame. [Full character descriptor — heritage, build, skin tone and finish, hair (color, length, texture), face register (jaw, chin, lips, cheekbones, brow shape), eye shape and color, all identity markers being locked: piercings with exact position and metal, scars with placement and size, beauty marks with placement, default makeup register]. She wears [a plain black thin-strap camisole / he wears a plain black ribbed tank], no jewelry, no logos, no graphics. Body squared to camera, head level, neutral relaxed expression, eyes to camera, lips closed and relaxed, subtle controlled energy.
|
||||||
|
[+ die identische Mid-Gray-Rembrandt-Schlusspassage wie in Step 0.A]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Why two steps for Soul Cinema:** Soul Cinema is faster and looser than Banana Pro on faces but holds less fidelity. The two-step flow uses Soul Cinema for exploration (cheap variations on the face register) and Banana Pro for the lock. This is the slowest path of the three options — only use it when iteration is more valuable than speed.
|
||||||
|
|
||||||
|
**What Mode 0 is NOT for:** refining an existing character (skip to Mode 1) · outfit design (Mode 1) · multi-angle sheets (Mode 2, only AFTER Mode 0 + Mode 1). Mode 0 is one-and-done per character.
|
||||||
|
|
||||||
|
## MODE 1A — SINGLE-IMAGE CHARACTER OUTFIT, BANANA PRO PATH
|
||||||
|
|
||||||
|
**When to use:** First image of any character/outfit pairing when the user picks Banana Pro. Best for relatively simple outfits where full prompt control gets us there in one clean shot.
|
||||||
|
|
||||||
|
**Goal:** Single character, face clearly readable, full styling locked head-to-toe, environment minimal so the character is the only subject.
|
||||||
|
|
||||||
|
**Frame and composition:** Subject centered, weight shifted onto one hip in the cocked-hip model stance, body angled 15–30° from camera, chin slightly tucked or level, eyes to camera or slightly off-camera. Default is full-body. Do not write aspect ratios into the prompt. Background: mid-gray seamless studio (locked default; white on explicit request → full cinema stack close). Lighting: soft soft cinematic key from camera-left or camera-right, very diffused, gentle wrap, no harsh shadows, no rim light, no hair light, no kicker.
|
||||||
|
|
||||||
|
**Default expression:** Model face-card neutral, subtle controlled, slight closed-lip smirk at most. Never teeth-showing smile unless the user specifically requests it.
|
||||||
|
|
||||||
|
**Canonical Mode 1A prompt structure:**
|
||||||
|
|
||||||
|
```
|
||||||
|
[Visual descriptor of the character — hair, makeup, full wardrobe head-to-toe, jewelry, body markers, all extracted from references or locked from the development phase]. [Pose direction — body angle, weight distribution, hand position, expression].
|
||||||
|
Mid-gray seamless studio background — even neutral mid-gray, no seam line, no gradient, no falloff to black or white. Relight from scratch overriding any reference lighting: one broad diffused source from camera-[left/right] and slightly above, gentle wrap onto the figure, no harsh shadows, no rim light, no hair light, no kicker, only the gentlest lifted shadow on the off-light side. Skin and fabric read matte and velvety in a low-contrast milky look, no shine, no oily T-zone. Skin renders at its true natural skin tone and the outfit at its true natural color, warmth preserved and natural against the neutral gray, never pale or washed-out or cool-shifted by the background. Real peach fuzz at the jaw and hairline, real fine even pore texture, subsurface scattering reading as semi-translucent biology, real fabric weave and drape, never plastic, never waxy, never harsh. Photographed on a 50mm prime at a wide aperture, natural round bokeh, even sharpness, soft natural film grain. Photographed not generated. [Framing — full body / waist-up / head-to-shoulders].
|
||||||
|
```
|
||||||
|
|
||||||
|
**Variation strategy when building multiple base references:** keep the mid-gray backdrop locked and vary ONE parameter per shot — Pose (cocked-hip front → angled three-quarter → seated → side profile → back-to-camera over-shoulder) · Framing (full body → waist-up → head-to-shoulders) · Expression (neutral → smirk → eyes-closed → looking off-frame) · Lighting direction (key from L → R → top → backlit). Don't vary face, skin, or core identity markers.
|
||||||
|
|
||||||
|
## MODE 1B — SINGLE-IMAGE CHARACTER OUTFIT, SOUL CINEMA PATH
|
||||||
|
|
||||||
|
**When to use:** when the user picks Soul Cinema — designing a custom fit separately from casting. Two-step flow, do not skip Step 1B.1.
|
||||||
|
|
||||||
|
### Step 1B.1 — Generate the outfit on a neutral model
|
||||||
|
|
||||||
|
Model spec (locked): slim model build, refined proportions · normal hair (medium-length straight or slight wave for women, short clean cut for men, medium brown by default) · normal model face, neutral natural makeup if a woman, blank neutral expression · straight-on stance, weight even, arms relaxed, body squared, eyes to camera · gender matched to the outfit.
|
||||||
|
|
||||||
|
```
|
||||||
|
A slim [woman / man] standing straight-on to camera in a relaxed neutral stance, weight evenly distributed across both feet, arms hanging relaxed at the sides, shoulders level and relaxed, body squared to the camera, head level. Medium-length [natural medium brown hair, simple straight or slight natural wave, parted naturally / short clean haircut, natural medium brown color]. Clean even features, neutral natural skin tone, [light natural makeup with skin-tint finish, soft groomed brows, neutral lip / no makeup, naturally groomed brows], neutral blank model expression, eyes directly to camera, lips closed and relaxed. Slim model build with refined proportions. The figure wears [full outfit description here — every garment top to bottom with fabric, color, fit, structural details, layering, hem positions, footwear, jewelry, accessories].
|
||||||
|
Mid-gray seamless studio background, even neutral mid-gray, no shadow falloff to black or white, no visible seam line, perfectly even backdrop. Soft soft natural light from camera-[left/right], very diffused, gentle wrap onto the figure, no harsh shadows, no dramatic rim light, no kicker, no hair light — only the gentlest lifted shadow on the off-light side. Skin and fabric read matte and slightly diffused, clean and even, the outfit fully readable and rendering at its true natural color against the neutral gray, never cool-shifted or washed-out by the background. Full body framing from head to just below the footwear.
|
||||||
|
Real fabric texture with visible weave detail, real weight, real drape, visible texture variation across the surface. Jewelry with real metal surface detail. Real human skin with natural pore texture. Fine cinema grain, soft lens vignette, natural color grade. Photographic, not rendered.
|
||||||
|
```
|
||||||
|
|
||||||
|
Lighting is intentionally soft soft from a single side — no full theatrical cinema stack at this stage. The outfit is the only subject.
|
||||||
|
|
||||||
|
### Step 1B.2 — Composite the outfit onto the locked character
|
||||||
|
|
||||||
|
Reference Image 1 = the locked character canonical. Reference Image 2 = the outfit reference from Step 1B.1.
|
||||||
|
|
||||||
|
```
|
||||||
|
Place the face and body from reference image 1 onto the outfit from reference image 2. Mid-gray seamless studio background, even neutral mid-gray, skin and outfit at their true natural tone. Soft studio lighting.
|
||||||
|
```
|
||||||
|
|
||||||
|
That's it. Do not add styling description, character description, cinema stack, or framing instructions unless requested. **When to push back to Mode 1A:** single-shot build or extreme stylistic control over how the outfit reads on the specific body.
|
||||||
|
|
||||||
|
## MODE 2 — 6-PANEL CHARACTER SHEET (SINGLE 16:9 FRAME)
|
||||||
|
|
||||||
|
Only after a single-image base reference is approved. Never six separate prompts — one prompt, one image, six panels, 3×2 grid, thin clean white gutters:
|
||||||
|
|
||||||
|
1. Top-left — Full body front · 2. Top-center — Side profile close headshot (left side) · 3. Top-right — Full body back · 4. Bottom-left — Side profile close headshot (right side) · 5. Bottom-center — Front face close headshot · 6. Bottom-right — Detail shot (nails / jewelry / piercing / tattoo / held prop — user picks).
|
||||||
|
|
||||||
|
**Canonical Mode 2 prompt structure:**
|
||||||
|
|
||||||
|
```
|
||||||
|
A 6-panel character reference sheet arranged as a 3-column by 2-row grid in a single horizontal frame, separated by thin clean white gutters between panels. Each panel shows the same single character — [full visual descriptor of the character including build, face, hair, makeup, full wardrobe head-to-toe, all accessories, jewelry, body markers, held props].
|
||||||
|
Panel 1 (top-left): Full body front — [stance description, framing, what's readable].
|
||||||
|
Panel 2 (top-center): Side profile close headshot, left side — [tight crop from collarbone up, character's left profile facing screen-right, hair and ear and jaw geometry visible].
|
||||||
|
Panel 3 (top-right): Full body back — [stance, what's visible from behind].
|
||||||
|
Panel 4 (bottom-left): Side profile close headshot, right side — [tight crop from collarbone up, character's right profile facing screen-left, mirror of Panel 2].
|
||||||
|
Panel 5 (bottom-center): Front face close headshot — [tight crop from collarbone up, body squared to camera, face filling the frame, eyes to camera].
|
||||||
|
Panel 6 (bottom-right): Detail shot — [the locked detail close-up: nails / specific jewelry piece / piercing / tattoo / held prop, filling the panel cleanly].
|
||||||
|
Mid-gray seamless studio backdrop applied uniformly across all six panels — even neutral mid-gray, no seam line, no gradient. Relight from scratch overriding any reference lighting, applied uniformly across all six panels: one broad diffused source from camera-left and slightly above, gentle wrap, no harsh shadows, no rim light, no hair light, no kicker. Skin and fabric read matte and velvety in a low-contrast milky look, rendering at their true natural skin tone and color against the neutral gray, warmth preserved and natural, never cool-shifted or washed-out by the background. Sharp focus across every panel. Real fine even pore texture, peach fuzz at the hairline, subsurface scattering, real fabric weave, soft natural film grain, photographed not generated. Identical character identity locked across all six panels — same face, same skin, same hair, same wardrobe, same accessories, same proportions in every cell.
|
||||||
|
```
|
||||||
|
|
||||||
|
Critical rules: identity described once in the opening paragraph; each panel only describes what's different (stance, angle, framing, focus); aspect ratio in the UI, never in the prompt; lighting and backdrop uniform; every panel gets its explicit position label.
|
||||||
|
|
||||||
|
## MODE 3 — CINEMATIC SCENE PLATE
|
||||||
|
|
||||||
|
**When to use:** only when the user asks for a scene/environment/plate/moment. Two flavors: **3A** character-in-environment (feeds Seedance video), **3B** pure environment plate.
|
||||||
|
|
||||||
|
**Camera grammar — five cinema modes paired to scene type:** M1 Narrative (real-world dramatic) · M2 Studio/Editorial · M3 Action/Combat · M4 Performance/Concert · M5 Atmospheric/Empty. The mode carries lens character, filtration look, film-stock rendition, grain, grade, color cast — described as the visual look, never as brand names. Woven into the closing camera-spec paragraph, e.g. "Captured with a wide-latitude cinema look and a vintage 55mm-equivalent 2x anamorphic character at a wide aperture — oval bokeh, gentle horizontal squeeze, soft frame-edge falloff, a light diffusion bloom lifting highlights into a soft halation, color-negative daylight film rendition with fine 35mm grain, in an M1 cinematic narrative register."
|
||||||
|
|
||||||
|
### The silent 6-block mental checklist (pre-composition only — never labeled blocks in output)
|
||||||
|
|
||||||
|
1. **Shot DNA** — camera position, what it looks at, framing register, mood. 2. **Subject behavior + spatial placement** — positional prose, not coordinates. 3. **Visible detail (resolution-aware)**. 4. **World** — ambience, not architecture; references carry geometry. 5. **Light and atmosphere**. 6. **Camera spec + finish** — continuous prose ending in the closing realism clause.
|
||||||
|
|
||||||
|
### Resolution-aware detail rule (locked)
|
||||||
|
|
||||||
|
Describe what the camera at this position can physically see, not what's "true" about the subject. Three silent diagnostics: (1) At this distance, would a real cinema lens resolve this detail? (2) At this motion blur level, would it read? (3) At this lighting register, would it be visible? If no → drop it. Examples: a car at 200 feet at 120 mph at dawn = silhouette + color blocks + headlights + motion trails (no decals, no badges); a person at 50 yards = silhouette + hair color + wardrobe blocks + posture; one-practical night face = face shape + eye glints + key wardrobe catching light. The same subjects in tight static close range get their detail described. Detail is earned by camera proximity, lens length, motion stillness, and lighting intensity.
|
||||||
|
|
||||||
|
### X/Y coordinate system (mental composition tool — never output notation)
|
||||||
|
|
||||||
|
Internal planning grid: X 0 % left → 100 % right, Y 0 % top → 100 % bottom, subject as bounding-box range. Rule-of-thirds anchors (33/67), placement library (hero subject left/right third, two-shot facing, wide environmental with hero on lower-right third, close-up with eye line on upper third, three-quarter portrait, horizon on upper/lower third, vehicle-in-motion with lead room, aerial light source entering top edge, architectural symmetry). Coordinates get translated into positional prose: "centered in the frame" · "in the left half of the frame" · "in the right portion of the frame" · "anchored to the lower-left third" · "the horizon line sitting at the upper third" · "anchored on the left third" · "in the deeper right background / positioned camera-right".
|
||||||
|
|
||||||
|
### The locked tag block (deprecated fallback only)
|
||||||
|
|
||||||
|
```
|
||||||
|
[Cinema mode tag — M1 Narrative / M2 Studio / M3 Action / M4 Performance / M5 Atmospheric]. Atmospheric volumetric haze. Real volumetric light physics. Gentle filmic highlight roll-off. Lifted blacks. Theatrical 35mm grain. Photographed not generated.
|
||||||
|
```
|
||||||
|
|
||||||
|
Only used when the user explicitly requests a stripped-down lean Mode 3 prompt. Default is the cinema-prose closing paragraph.
|
||||||
|
|
||||||
|
### The cinema-prose register (locked, non-negotiable)
|
||||||
|
|
||||||
|
Mode 3 prompts are written like a DP describing a real frame, not like a spec sheet. No labeled headers, no coordinate notation, no CRITICAL LIGHTING RULES blocks, no explicit negations, no architectural enumeration. The voice is cinematic anamorphic prose — confident, declarative, observational.
|
||||||
|
|
||||||
|
**The five-paragraph prose structure (locked):**
|
||||||
|
|
||||||
|
1. **Opening shot description** — one long sentence: medium, framing register, subject at high level, camera position and angle in prose, mood/intent.
|
||||||
|
2. **Character block** — identity markers from the attached reference written as visible facts; pose, attention, held props woven in.
|
||||||
|
3. **World/environment block** — location as ambience and atmosphere; anchored to the attached reference; background subjects in positional language.
|
||||||
|
4. **Subject anchor block** — the focal anchor (TV broadcast, second car, horizon) gets its own paragraph; folds into 3 if none.
|
||||||
|
5. **Camera spec + finish** — full cinema look in one continuous paragraph, plain-language look terms, never brand names, ending in the closing realism clause ("Real photographic frame captured on a real cinema camera … — no CGI, no rendered look, no digital cleanliness, no plastic surfaces, no AI smoothness, no skin smoothing, no glow, no halation bloom that reads as artificial, no glossy highlights"). The closing realism clause is mandatory.
|
||||||
|
|
||||||
|
**Key writing rules:** no labeled blocks in output · no coordinates in the body · no CRITICAL/IMPORTANT/MUST rules · no explicit negations as instructions (write what IS there; negations only in the closing realism clause) · references do the geometry work ("carrying identically from the attached world reference") · references do the identity work · the prompt narrates THE MOMENT · closing realism clause non-negotiable · cinema mode invoked by describing the look, M-tag as brief identifier in prose · no aspect ratios.
|
||||||
|
|
||||||
|
**Canonical Mode 3 reference example:**
|
||||||
|
|
||||||
|
```
|
||||||
|
A cinematic anamorphic still photograph captured handheld on a real cinema set — a low-angle medium hero composition of a woman standing alone at the edge of an empty rooftop at dusk, the camera positioned slightly below her eye line in a waist-up framing anchored to the left third of the frame, the deepening dusk sky filling the upper two-thirds of the frame behind her, the city skyline reading in soft silhouette across the lower third of the background, the composition holding a quiet observational stillness.
|
||||||
|
The character carrying identically from the attached character reference — her hair, skin, makeup, and identity locked from the reference. She wears the wardrobe carrying identically from the attached wardrobe reference, the fabric reading natural across her shoulders and upper torso. Her body is angled three-quarters toward camera, her weight settled on her back foot, her left hand resting loosely at her side, her right hand at her hip. Her gaze is locked across the rooftop toward the horizon screen-right, her expression neutral and held, her shoulders relaxed but settled.
|
||||||
|
The rooftop beyond her is the location carrying from the attached environment plate — weathered concrete edge, rusted railing in the foreground softened by shallow depth of field, the city skyline beyond reading as silhouette layers stacked into atmospheric haze, distant building lights coming on one by one as dusk falls. Light atmospheric haze suspended through the deeper space giving the air real physical body, the horizon glow warm magenta-orange transitioning into deep blue overhead. Practical warm light from off-frame at camera-right catches the right side of her face and shoulder with restrained natural rim, the cool ambient dusk light wrapping faintly around her left side where the warm and cool temperatures meet.
|
||||||
|
The city skyline reads as the visual anchor of the deeper frame — building silhouettes layered front-to-back with progressive atmospheric desaturation, the warm horizon glow visible between the structures, scattered building lights warm and small in the deep distance, a faint aircraft beacon blinking once at the upper-right edge of the frame, the rest of the sky held in deep cool blue with the first stars just visible at the upper edge.
|
||||||
|
Captured with a wide-latitude cinema look and a vintage 55mm-equivalent 2x anamorphic character at a wide aperture, a light diffusion bloom softening the highlights, color-negative daylight film rendition pushed slightly, in an M1 cinematic narrative register. Real anamorphic optical character with oval bokeh on the deeper city elements, organic handheld operator breath, subtle frame-edge falloff, a faint horizontal streak flare on the brightest horizon highlight. Theatrical fine 35mm film grain across the entire frame — skin, fabric, concrete, sky, haze. Contemporary teal-amber cinema grade with the warm horizon glow on her right side meeting the cool dusk wash on her left, shadows lifted gently into deep cool blue-grey never crushed, highlights rolled off softly never blown. Real photographic frame captured on a real cinema camera, real anamorphic lens, real fabric, real human subject, real concrete and haze — no CGI, no rendered look, no digital cleanliness, no plastic surfaces, no AI smoothness, no skin smoothing, no glow, no halation bloom that reads as artificial, no glossy highlights.
|
||||||
|
```
|
||||||
|
|
||||||
|
The old coordinate grammar (labeled blocks, X/Y notation, CRITICAL LIGHTING RULES, explicit negations, room enumeration) is **deprecated** — it made the model overcorrect and confuse spatial relationships.
|
||||||
|
|
||||||
|
## MODE 4 — GPT-2 DETAIL FACE SHOT (HIGGSFIELD GPT-2)
|
||||||
|
|
||||||
|
Only when the user explicitly asks for a chest-up portrait, face detail shot, or close-up where face/skin/eye fidelity matters most. Gate: confirm GPT-2 + credit cost first (once per conversation).
|
||||||
|
|
||||||
|
Framing: chest-up, shoulders-up, or face-only. Background: mid-gray seamless (default) or soft moody studio backdrop. Lighting: classical beauty lighting — soft key from slightly above and camera-left, soft fill at chest level from camera-right, subtle hair light behind, soft underlight bounce lifting the eye sockets.
|
||||||
|
|
||||||
|
```
|
||||||
|
[Visual descriptor of the character — hair, makeup, wardrobe visible in frame from the chest up, jewelry visible at collar and ears, eye color and detail, lip detail, skin finish]. [Pose direction — head angle, shoulder angle, expression register].
|
||||||
|
[Background — mid-gray seamless studio (locked default) OR specified moody backdrop]. Classical beauty lighting — soft key from slightly above and camera-left at 35 degrees, soft fill at chest level from camera-right, subtle hair light behind defining the crown, soft underlight bounce lifting the eye sockets. [Framing — chest-up portrait / shoulders-up / face-only forehead-to-collarbone].
|
||||||
|
Extreme face fidelity. Real skin texture with visible pores, fine peach fuzz catching light along the jawline and upper lip, subtle subsurface scattering on the nose bridge cheeks and ears, micro-expression detail in the eyes and mouth corners, individual lash detail, real moisture and reflection in the iris with visible iris pattern, real lip texture with subtle natural lip lines, hair rendered strand by strand at the hairline with visible baby hairs and flyaways, fabric weave visible at the collar and shoulder.
|
||||||
|
[The cinema stack].
|
||||||
|
```
|
||||||
|
|
||||||
|
Why GPT-2 for these shots: stronger read on micro-detail at face-and-shoulders range — pores, lash separation, iris pattern, lip texture, hairline strand definition.
|
||||||
|
|
||||||
|
## MODE 5 — OUTFIT REPLACEMENT (BANANA PRO TWO-REFERENCE SWAP)
|
||||||
|
|
||||||
|
Outfit and pose from one image applied to a different character. **@image1 = outfit reference** (wardrobe, styling, footwear, accessories, pose to keep). **@image2 = character reference** (face, bone structure, body type, skin tone, hair to apply). This order is fixed — reversing it breaks the swap.
|
||||||
|
|
||||||
|
**Canonical Mode 5 prompt (LOCKED — do not modify):**
|
||||||
|
|
||||||
|
```
|
||||||
|
Replace the character in @image1 with the character in @image2. Keep the outfit and pose from @image1 exactly. Match the face, bone structure, body type, skin tone, and hair from @image2. Clean mid-gray seamless studio background, even neutral mid-gray with no seam line, soft large-source studio lighting, skin and outfit rendering at their true natural tone against the neutral gray, natural film grain, full body framing.
|
||||||
|
```
|
||||||
|
|
||||||
|
Mode 5 does not use the cinema stack — the references carry the photographic register; extra language degrades the identity transfer. Background/lighting language locked. Character- and IP-agnostic, no per-project modifiers. For a different environment: run Mode 5 first, then Mode 3 on the output.
|
||||||
|
|
||||||
|
## UNIVERSAL PROMPT RULES (ALL MODES)
|
||||||
|
|
||||||
|
1. No character names in prompt output.
|
||||||
|
2. No real brand names in prompt output.
|
||||||
|
3. No `@image` tags or `<<<image_n>>>` placeholders (Ausnahme: Mode 5) — image attachment happens in the Higgsfield UI.
|
||||||
|
4. No internal production context — every prompt is standalone.
|
||||||
|
5. Pure visual description only — no meta-commentary, no emotional intent.
|
||||||
|
6. No teeth-showing smiles unless explicitly requested.
|
||||||
|
7. No negative prompt blocks — Higgsfield workflow doesn't use them.
|
||||||
|
8. Cinema stack baked in for Modes 0, 1, 2, 4, 5 (Mode 1B Step 1: lighter close; Mode 5: locked lean prompt).
|
||||||
|
9. Mode 3 uses the cinema-prose closing paragraph in place of the cinema stack AND the locked tag block.
|
||||||
|
10. Single fenced code block on output.
|
||||||
|
11. Pre-prompt confirmation always — except minor iteration on an approved prompt.
|
||||||
|
12. No aspect ratios in prompt output — framing in plain language only.
|
||||||
|
|
||||||
|
## INVENTORY EXTRACTION CHECKLIST (run silently before composing)
|
||||||
|
|
||||||
|
Mode selected + rationale · every reference identified and listed · Mode-0/1/2/4/5-Voraussetzungen erfüllt (Face-Lock existiert, Base-Reference existiert, GPT-2 bestätigt, Referenz-Reihenfolge bestätigt) · every character described by visual markers only · Mode 3: ambience not architecture, cinema mode gewählt, positional prose, resolution-aware check, five-paragraph structure, closing realism clause · pose/body angle/expression chosen · no names/brands/context/meta · correct closing per mode · pre-prompt confirmation delivered. If anything needed is missing from the user input, ask before writing.
|
||||||
|
|
||||||
|
## WHEN THE USER ASKS FOR A PROMPT
|
||||||
|
|
||||||
|
The flow is always: confirm character → confirm what's about to be prompted → deliver the prompt in a fenced code block. Tool routing: Banana Pro/Nano Banana for Mode 0 Step 0.A, Step 0.2, Modes 1A, 2, 3, 5; GPT-2 for Step 0.B and Mode 4; Soul Cinema for Step 0.1 and Mode 1B. Multiple shots in one ask → each in its own code block, one pre-prompt confirmation for the batch.
|
||||||
266
docs/cinema-worldbuilder-pro-2.0.md
Normal file
266
docs/cinema-worldbuilder-pro-2.0.md
Normal file
@@ -0,0 +1,266 @@
|
|||||||
|
# Cinema Worldbuilder Pro 2.0 — Seedance Director
|
||||||
|
|
||||||
|
> **Quelldokument.** Original-Skill (Higgsfield/Seedance-Workflow, interaktiv). Dient als
|
||||||
|
> Quelle für [prompts/P08-video-kern.md](../prompts/P08-video-kern.md) und
|
||||||
|
> [prompts/P08-adapter-seedance.md](../prompts/P08-adapter-seedance.md) — die
|
||||||
|
> Pipeline-Destillate ohne die interaktiven Teile. Wird NICHT geseedet.
|
||||||
|
> Formatierung gegenüber der Chat-Übergabe minimal bereinigt, Inhalt vollständig.
|
||||||
|
|
||||||
|
The locked cinematography grammar for Seedance video prompts. This skill is mode-aware, reference-aware, composition-aware, and audio-aware. It reads what the user gives you, picks the right cinema mode, extracts wardrobe and identity from reference images by visual description, maps the frame, locks every character to a screen position and state, choreographs the motion, fixes the closing composition, and outputs a production-ready Seedance prompt with diegetic audio only.
|
||||||
|
|
||||||
|
Pro 2.0 is built around density discipline: shorter prompts render better than longer ones. Every block does work. Nothing is decorative. The Camera Capture spec is one trimmed line at the bottom — never doubled. The Subject Lock trusts the reference image to carry wardrobe and identity, naming only what the model cannot read from the image itself (pose, gaze, state, contact points, what stays unchanged).
|
||||||
|
|
||||||
|
## CORE PHILOSOPHY
|
||||||
|
|
||||||
|
No plastic. No commercial gloss. No LED-panel-rendered-on-a-soundstage energy. No Instagram-ad sharpness.
|
||||||
|
|
||||||
|
Every frame should feel captured on a camera that has lived a little — film-emulated, filtered, slightly imperfect, analog warmth in the highlights, controlled blacks that aren't crushed. The grade is editorial, not commercial. The glass has character. The shadows hold detail. Real fabric, real skin, real sweat, real haze, real grain.
|
||||||
|
|
||||||
|
Five modes share a wide-latitude cinema capture look and either a vintage 2x anamorphic character or a clean spherical character. The differences across the modes are in **movement, diffusion, grade, palette, and texture** — not in capture register or lens family.
|
||||||
|
|
||||||
|
A great prompt is not a beautiful sentence. It is a production document. Seedance follows physical, spatial, and cinematographic logic far better than abstract poetry. Every shot answers: who is in the frame, where exactly they sit, what state they hold, what moves, what stays locked, how the camera operates, and what the final frame must look like.
|
||||||
|
|
||||||
|
**Density rule.** Target prompt length is 280–400 words for single-shot scenes. Multi-shot sequences may run longer but never over 600. Every word should do work. When in doubt, trust the reference image to carry visual information and cut the redundant description.
|
||||||
|
|
||||||
|
## HOW TO USE THIS SKILL
|
||||||
|
|
||||||
|
**Step 1 — Upload reference material.** Character images, environment plates, mood references, wardrobe shots. Purely environmental or invented-from-scratch scenes need no images.
|
||||||
|
|
||||||
|
**Step 2 — Describe the scene.** Who is in the frame, what they're doing, where it's set, what's happening, how long the shot should run. The skill picks the cinema mode automatically (or the user names it).
|
||||||
|
|
||||||
|
**Step 3 — Confirm the pre-prompt summary.** Bulleted check: references (first), mode, scene, characters, frame map, camera, runtime (last).
|
||||||
|
|
||||||
|
**Step 4 — Receive the three-part delivery.** (a) numbered bulleted list of reference images to attach in order (max 9 — Seedance hard cap), (b) bolded English title line with runtime, (c) single fenced English code block with discrete labeled blocks in exact order — Scene & Mood → Frame Map → Subject Lock(s) → Cross-Frame Rules → Movement → Last Frame → World Plate → Sound Bed → Capture Realism → Camera Capture — with inline `@image1`–`@image9` tags matching the bullet list.
|
||||||
|
|
||||||
|
**Step 5 — Run it in Higgsfield.** Attach the references in the exact order listed, paste the code block. The `@imageN` tags are functional Seedance syntax.
|
||||||
|
|
||||||
|
## SESSION OPENER — CHARACTER GATE
|
||||||
|
|
||||||
|
First Seedance prompt of a session, ask once: "Any recurring characters in this batch? If so, are they already built (reference images locked) or do we need to develop them first?" — Yes/built → reference upload, study and lock, mirror back the spec. Yes/needs developing → kick over to banana-pro-director's character development flow first. No → skip. Once asked, do not ask again in the session.
|
||||||
|
|
||||||
|
## PRE-PROMPT CONFIRMATION RULE
|
||||||
|
|
||||||
|
Every NEW scene gets a pre-prompt summary before the full prompt:
|
||||||
|
|
||||||
|
```
|
||||||
|
Pre-prompt check:
|
||||||
|
- **References attached:** [every reference by short visual descriptor; or "none — pure text composition."]
|
||||||
|
- **Mode:** [M1 Narrative / M2 Studio / M3 Action / M4 Performance / M5 Atmospheric]
|
||||||
|
- **Scene:** [one-line scene description]
|
||||||
|
- **Characters:** [who's in frame, abbreviated by visual marker; or "none / environment plate"]
|
||||||
|
- **Frame Map:** [one-line compositional read]
|
||||||
|
- **Camera:** [lens length, key movement]
|
||||||
|
- **Runtime:** [Xs, single shot, OR Xs, N-shot sequence]
|
||||||
|
|
||||||
|
Sound good?
|
||||||
|
```
|
||||||
|
|
||||||
|
References first (confirms every upload is being used), runtime last (the most important spec sits right above "Sound good?"). Skip only on: iteration of a just-delivered prompt, pre-confirmed batches, or explicit "skip the confirm". **Runtime: always ask, never assume a default.**
|
||||||
|
|
||||||
|
## THREE-PART DELIVERY FORMAT (LOCKED)
|
||||||
|
|
||||||
|
1. Numbered bulleted reference list (max 9).
|
||||||
|
2. Bolded English title line with runtime, e.g. `**Seedance prompt — 12s**`.
|
||||||
|
3. English code block with the ten labeled blocks and inline `@imageN` tags (bullet 1 = `@image1` …).
|
||||||
|
|
||||||
|
**Block order inside the code block (every prompt):**
|
||||||
|
|
||||||
|
```
|
||||||
|
Scene & Mood: [one or two sentences — what the moment IS, dramatically]
|
||||||
|
|
||||||
|
Frame Map: [where each subject sits — thirds, depth, x% where helpful, negative space; per-shot for sequences]
|
||||||
|
|
||||||
|
Subject Lock — @imageN: [per character — identity anchor + body orientation + pose + state + gaze + contact points + lock-down line. Trust the reference for wardrobe; only re-describe what the image can't carry]
|
||||||
|
|
||||||
|
Cross-Frame Rules: [multi-character: never swap, never cross center, never change depth, distance and screen sides held. Multi-shot: what carries across the cut]
|
||||||
|
|
||||||
|
Movement: [character motion + micro-motion + environmental motion across the runtime, flowing paragraph with per-beat timestamps]
|
||||||
|
|
||||||
|
Last Frame: [exact closing composition + on-screen text suppression line]
|
||||||
|
|
||||||
|
World Plate: [location, time, weather, set dressing, atmosphere — anchored to @imageN if a plate is attached]
|
||||||
|
|
||||||
|
Sound Bed: [diegetic only — specific sounds, no music, no lyrics, no score]
|
||||||
|
|
||||||
|
Capture Realism: [anti-plastic/anti-contrast block — depth via suspended atmosphere, moisture-without-shine if wet, per-zone specular kill, contrast curve three ways]
|
||||||
|
|
||||||
|
Camera Capture: [single trimmed paragraph — body, lens, filter, movement, stock, grade, frame rate, runtime]
|
||||||
|
```
|
||||||
|
|
||||||
|
## OUTPUT LANGUAGE (LOCKED)
|
||||||
|
|
||||||
|
English only inside the code block. Aesthetic descriptors in plain-language English (wide-latitude cinema capture, vintage 2x anamorphic character, soft diffusion bloom, color-negative film rendition, fine 35mm grain) — never brand names or model numbers. Numerals for real optical properties (mm, 24fps, 180° shutter). No Chinese mode, no bilingual mode.
|
||||||
|
|
||||||
|
## UNIVERSAL PROMPT RULES (ALL MODES)
|
||||||
|
|
||||||
|
1. Pre-prompt confirmation on every new scene (references FIRST, runtime LAST).
|
||||||
|
2. Three-part delivery format, in order.
|
||||||
|
3. `@imageN` numbering matches the bullet list exactly.
|
||||||
|
4. Every listed reference appears at least once as an `@imageN` tag.
|
||||||
|
5. Runtime baked into the closing Camera Capture line; always ask; title and Camera Capture must match.
|
||||||
|
6. Per-shot timing inline in Movement for any multi-cut sequence.
|
||||||
|
7. **Discrete labeled blocks, exact order, every prompt — HARD LOCK:** Scene & Mood → Frame Map → Subject Lock(s) → Cross-Frame Rules → Movement → Last Frame → World Plate → Sound Bed → Capture Realism → Camera Capture. No block omitted, reordered, merged, renamed, or replaced with flowing prose. Conditional content only INSIDE blocks (IF-WET-Klausel, Haut-Satz bei M5); leere Blöcke werden gekürzt, nie weggelassen.
|
||||||
|
8. One Subject Lock block per character — never jammed into one paragraph.
|
||||||
|
9. One Camera Capture line at the bottom — never doubled; the only camera/grade/stock language in the prompt.
|
||||||
|
10. No character names in prompt output.
|
||||||
|
11. No real brand names ("white low-slung mid-engine sports car").
|
||||||
|
12. No platform/tool names (never "Higgsfield," "Seedance," "Banana Pro," "Soul Cinema") inside the prompt text.
|
||||||
|
13. No internal production context — every prompt standalone.
|
||||||
|
14. Pure visual description only — no meta-commentary.
|
||||||
|
15. Diegetic audio only — no music, no lyrics, no song references.
|
||||||
|
16. Energy over position in Scene & Mood; Frame Map handles geometry.
|
||||||
|
17. Cut triggers: "Hard cut to," "Smash cut to," "Match cut on."
|
||||||
|
18. Age-blind — describe by role, hair, wardrobe, identity markers.
|
||||||
|
19. No on-screen text by default; every Last Frame closes with "No on-screen text, no captions, no signage typography, no rendered text in the frame." (skip only when text is explicitly requested).
|
||||||
|
20. Positive locks over negative prohibitions ("no drifting" → "boots stay planted on the same ground marks").
|
||||||
|
21. One main idea per shot — one dominant action, one camera strategy, one lighting motivation; split otherwise.
|
||||||
|
22. Trust the reference image for wardrobe — only state-changes the image can't carry (damp, torn, dusty).
|
||||||
|
23. **Canonical reference always attached, never substituted by the plate (HARD LOCK).** Every named subject gets its canonical reference as its own `@imageN` slot, even when visible in the environment plate. Plate carries the world; canonical reference carries identity. Subject Locks anchor to canonical tags; World Plate anchors to the plate tag. No exceptions — this prevents identity drift.
|
||||||
|
|
||||||
|
## READING REFERENCE IMAGES
|
||||||
|
|
||||||
|
Extract everything visible by visual description only — never names, never invented details. Per character: hair (every nuance), makeup, wardrobe (every garment, generically for brands), jewelry & accessories, body markers (only if visible), pose and energy. Per environment: location, time of day and weather, set dressing, color palette. The extracted reading is for understanding and the pre-prompt check; the prompt body trusts the reference and restates only what the image cannot carry (lock-down line: "face, hair, wardrobe, and silhouette identical throughout").
|
||||||
|
|
||||||
|
## FRAME MAP
|
||||||
|
|
||||||
|
Anchors every subject in screen space before motion. 2D screen space: horizontal (thirds or x%), vertical (thirds or y%), depth (foreground/midground/background), frame occupancy (close-up … extreme close-up, or % of frame height), negative space (what stays empty, where, filled with what).
|
||||||
|
|
||||||
|
> Frame Map: @image1 anchored in the left third, x=30%, foreground, medium shot from waist up, occupying 55% of frame height. The right two-thirds hold wet street and distant neon signage as negative space.
|
||||||
|
|
||||||
|
> Frame Map: @image1 in the left third, x=28%, foreground. @image2 in the right third, x=72%, midground, slightly deeper. The center holds open as tense negative space between them. Neither crosses the central vertical axis.
|
||||||
|
|
||||||
|
> Frame Map: Shot 1 (0–6s) — wide two-shot. @image1 in the left third, x=32%, foreground, bent at the waist. @image2 in the right third, x=68%, midground, leaning against @image3. Shot 2 (6–10s) — low-angle close-up at hip height looking up at the side window, framed tight on @image1's reflection in the wet glass.
|
||||||
|
|
||||||
|
Skip percentages for clear classical compositions (centered single, OTS, profile two-shot, symmetrical wide); coordinates earn their place when the composition is asymmetric, tightly blocked, or drift would break the shot.
|
||||||
|
|
||||||
|
## SUBJECT LOCK
|
||||||
|
|
||||||
|
Per character: identity anchor (`@imageN`) · body orientation · pose · state (emotional register via what body and face physically do — never abstract feelings) · expression (lips, eyes, brow, jaw) · gaze direction · contact points (feet on which surface, hand on which object) · state-change details the image can't carry (damp, dirty, torn, wet, dusty, bloodied) · lock-down line ("face, hair, wardrobe, and silhouette identical throughout").
|
||||||
|
|
||||||
|
> Subject Lock — @image1: Face, hair, oxblood corset, and silhouette identical throughout. Ponytail damp from the drizzle, fabric darker where rain has soaked in. Bent at the waist, torso angled toward the side window of @image3, both hands raised to her ponytail at the crown, fingers smoothing strands. Body squared to the car, weight even. Gaze locked on her own reflection in the wet glass.
|
||||||
|
|
||||||
|
Multiple characters → one discrete Subject Lock block each, never one paragraph.
|
||||||
|
|
||||||
|
## CROSS-FRAME RULES
|
||||||
|
|
||||||
|
For 2+ characters: no swap · no center crossing (unless an action demands it — then state the crossing with timing) · no depth change · distance consistency · screen sides held · eyelines (who looks at whom, holds or breaks) · carry-across-the-cut for sequences.
|
||||||
|
|
||||||
|
> Cross-Frame Rules: @image1 and @image2 never swap positions, never cross center, never change depth. Distance, screen sides, eyelines, costumes, and silhouettes stay consistent across the full runtime.
|
||||||
|
|
||||||
|
Crossing example: "At 4 seconds, @image1 steps across the central axis from the left third into the center. After 5 seconds, the new blocking holds."
|
||||||
|
|
||||||
|
## MOVEMENT
|
||||||
|
|
||||||
|
Four layers, in this order, in one flowing paragraph: 1. character motion (with per-beat timestamps) · 2. micro-motion (breath, hair, fabric, jewelry) · 3. environmental motion (rain, smoke, dust, traffic, wind) · 4. camera motion (usually omitted — Camera Capture handles it).
|
||||||
|
|
||||||
|
> Movement: She takes one slow controlled step from the curb to the street across the first two seconds, then holds for the remaining eight. Ponytail catching subtle wind drift, parachute pants fabric rustling on the step, breath visible in the cold air on a controlled exhale, fingers flexing once inside her front pockets. Light cold rain falling at moderate density, neon reflections shimmering on the wet asphalt, distant taxi headlights moving slowly through the right midground, faint steam rising from a manhole grate behind her.
|
||||||
|
|
||||||
|
Critical rule: never tangle the layers; each named explicitly — "nothing else moves in the frame" is a directive, absence is not.
|
||||||
|
|
||||||
|
## LAST FRAME
|
||||||
|
|
||||||
|
Mandatory closing block: where each character sits at the close · final pose/state/gaze · what the camera shows in focus · negative space at the close · the visual punctuation · the on-screen text suppression line.
|
||||||
|
|
||||||
|
> Last Frame: Hold on her in the left third, eyes still tracking the now-passed taxi offscreen right, ponytail settling, rain visible on her shoulders, the center of the frame filled with empty wet street and reflected neon, taxi taillights fading at the right edge. No on-screen text, no captions, no signage typography, no rendered text in the frame.
|
||||||
|
|
||||||
|
## WORLD PLATE
|
||||||
|
|
||||||
|
Location (anchored to @imageN if a plate is attached) · time of day and weather · set dressing · color palette · atmospheric quality (haze density, particles, weather intensity).
|
||||||
|
|
||||||
|
> World Plate: Anchored to @image4 — cliffside overlook with low grass and exposed rock at the edge, the drop falling away behind @image3, dusk sky dropping from cool blue at top into deep magenta and warm tungsten residue at the horizon, distant clouds, light atmospheric haze. @image3 parked perpendicular to the cliff edge, paint slick with rain, side windows wet, faint mist off the warm hood.
|
||||||
|
|
||||||
|
## SOUND BED
|
||||||
|
|
||||||
|
Only what the scene physically produces. **Allowed:** footsteps (with surface), fabric movement, breath, body sounds, object sounds, environmental ambient, mech/sci-fi diegetic, crowd diegetic, stage diegetic, weather. **Never:** song/artist/album names, lyrics, "music plays / soundtrack swells", score descriptors, genre cues.
|
||||||
|
|
||||||
|
**Audio modes:** Mode 1 (default) — diegetic with SFX and ambient: `Sound Bed: Diegetic only — [sounds], no music, no dialogue except what is physically spoken in frame.` · Mode 2 — silent capture (only when the user explicitly adds music in post AND wants silence): `Sound Bed: NONE — fully silent capture.` · Mode 3 — diegetic, no music explicitly.
|
||||||
|
|
||||||
|
## CAPTURE REALISM BLOCK (LOCKED — THE REAL-FOOTAGE ENGINE)
|
||||||
|
|
||||||
|
Camera Capture names the *gear*; this block names the *physics*. Second-to-last, ships on every prompt unless the user explicitly asks for a glossy/clean/commercial register. It attacks the three default AI-video failures: flat single-plane staging, glossy moisture/skin, over-rendered contrast.
|
||||||
|
|
||||||
|
**The four mechanics:**
|
||||||
|
|
||||||
|
1. **Depth via suspended atmosphere between planes** — default-on wherever there are planes to separate (M1/M3/M4/M5 always, M2 when depth exists); scale thin/light/heavy, never drop. Tie it to the actual planes of THIS shot.
|
||||||
|
2. **Moisture without shine** — only if the scene is wet/humid/sweaty: damp not beaded, wet not glossy, no specular hotspot. Bone-dry scene → skip entirely.
|
||||||
|
3. **Per-zone specular kill on skin + flattering ceiling** — name the zones: forehead, nose bridge, cheekbones, temples, chin, collarbones. Pair with biology cues (peach fuzz, soft pore texture, subsurface scattering, warmth preserved). Flattering ceiling locked: fine, soft, even — no acne, no blemishes, no scarring, no enlarged pores, no clinical macro-detail. Matte carries anti-plastic, fine-and-even carries flattering; tension resolves toward flattering.
|
||||||
|
4. **Contrast curve stated three ways** — (a) tonal curve: shadows lifted gently, highlights rolled off softly, nothing clipping or crushing; (b) specular removal: all speculars surgically removed, every pixel matte and diffuse; (c) grade: low-contrast, slightly desaturated, warmth preserved. Three statements hold; one gets overridden.
|
||||||
|
|
||||||
|
**Canonical block (tune every bracket to the scene):**
|
||||||
|
|
||||||
|
```
|
||||||
|
Capture Realism: [Foreground subject] sits inside real depth — [thin/light/heavy] atmosphere suspended in the air between camera, subject, and [the far background element], the background rendered softer, desaturated, and lower-contrast than the foreground so the figure sits within the air rather than pasted on a flat plane. [IF WET: Slight moisture has settled on every surface — damp matte hair, slight moisture on skin holding fully matte with no beading and no wet sheen, [wet ground with muted reflection / damp matte fabric / car paint damp but matte not showroom], moisture that mutes and deepens without a single specular hotspot.] Skin reads true cinematic matte — zero shine on forehead, nose bridge, cheekbones, temples, chin, and collarbones, real peach fuzz catching light at the jaw and hairline, real soft fine even pore texture, light absorbed like true subsurface scattering, warmth preserved and natural, slightly desaturated but never pale or washed-out or cool-shifted, never plastic, never doll-skin, never AI-rendered, and never harsh — no acne, no blemishes, no enlarged or rough pores, fine flattering texture that keeps the face looking good. Low-contrast curve — shadows lifted gently holding texture, highlights rolled off softly never clipping to white, nothing crushed to black. All specular highlights surgically removed from skin, hair, fabric, and surrounding surfaces, every pixel reading matte and diffuse. Slightly desaturated grade with warmth preserved.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Tuning:** dry → IF-WET-Satz löschen · no humans (M5) → Haut-Satz streichen, Matt-Logik auf Oberflächen · M2 glossy-editorial gewollt → reduzieren/skippen (einziger Mode mit bewusstem Specular) · Atmosphärendichte skaliert (thin Interior, light Exterior, heavy Pre-Dawn/Destroyed) · Block nennt NIE Gear/Hex/fps/Runtime — das ist Camera Capture. Lean positive; die Specular-Kill-/Anti-Plastik-Formeln sind die sanktionierte Negativ-Ausnahme.
|
||||||
|
|
||||||
|
## CAMERA CAPTURE
|
||||||
|
|
||||||
|
Single closing line: body, lens, filter, movement, stock, grade, frame rate, runtime — one trimmed paragraph. The ONLY camera/grade/stock language anywhere in the prompt. Default camera energy is handheld with breath, drift, and organic operator movement — locked-off tripod is OPT-IN only (explicit request or shot type that requires it).
|
||||||
|
|
||||||
|
## MODE-SELECT TABLE
|
||||||
|
|
||||||
|
| Mode | Use when scene is… | Lens | Movement | Grade |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| **M1 — Narrative** | Real-world dramatic, lived-in | Vintage 2x anamorphic, 40/55/75/100mm, oval bokeh, soft edge falloff | Handheld with operator breath | Color-negative daylight, fine 35mm grain, teal-amber |
|
||||||
|
| **M2 — Studio/Editorial** | White void, clean studio, fashion film, editorial | Clean spherical, 32/50/75/100mm, round bokeh, even sharpness | Locked tripod + optional slow push | Saturated editorial, warm-retained blacks, fine grain; intentional bloom on chrome/rhinestone |
|
||||||
|
| **M3 — Action/Combat** | Combat, chase, stunts, debris, smoke | Vintage 2x anamorphic, 40/55/75/100mm | Handheld and shaky throughout, no stabilized shots | Color-negative, heavier low-light grain, palette per scene, dusty haze |
|
||||||
|
| **M4 — Performance/Concert** | Stadium, stage, pit, lightstick crowd | Vintage 2x anamorphic, streak flares on stage lights | Mixed handheld pit-photographer + orbital, hard cuts | Color-negative, fine grain, desaturated cool with warm bloom, stage color cast |
|
||||||
|
| **M5 — Atmospheric/Empty** | No-humans plates, landscapes, weather, establishing | Vintage 2x anamorphic, 35→85mm push range | Locked-off or extremely slow push | Color-negative, fine grain, palette-driven (hex per scene) |
|
||||||
|
|
||||||
|
**Camera Capture lines per mode:**
|
||||||
|
|
||||||
|
M1: `Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, soft frame-edge falloff — light diffusion bloom softening highlights, handheld with natural operator breath, color-negative daylight film rendition with fine 35mm grain, teal-amber grade, shallow depth of field, 24fps 180° shutter, [XX] seconds.`
|
||||||
|
|
||||||
|
M1 multi-shot: `Camera Capture: Shot 1 — wide-latitude cinema capture, vintage 40mm 2x anamorphic character at a wide aperture, light diffusion bloom softening highlights, handheld with natural operator breath. Shot 2 — same capture register, 75mm anamorphic character at a wide aperture, low-angle handheld at hip height, tight operator breath. Color-negative daylight film rendition with fine 35mm grain, teal-amber grade, shallow depth of field, 24fps 180° shutter, [XX] seconds total.`
|
||||||
|
|
||||||
|
M2: `Camera Capture: wide-latitude cinema capture, clean spherical [XX]mm character at a wide aperture — natural round bokeh, even sharpness — mild diffusion bloom, locked tripod with optional slow push-in, saturated editorial grade, fine grain, warm-retained blacks, 24fps 180° shutter, [XX] seconds.` (+ bei Chrom/Strass: `intentional highlight bloom on reflective surfaces, blooming the speculars on chrome and rhinestone.`)
|
||||||
|
|
||||||
|
M3: `Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, soft edge falloff — light diffusion bloom softening highlights, handheld and shaky throughout with no stabilized shots, color-negative film rendition with heavier low-light grain, [palette descriptor] with dusty atmospheric haze, 24fps 180° shutter, [XX] seconds.` (Slow-Motion: `intercut 96fps high-speed slow-motion at the [moment] holding 180° shutter for natural motion blur.`)
|
||||||
|
|
||||||
|
M4: `Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, horizontal streak flares on stage lights — light diffusion bloom softening highlights, mixed handheld pit-photographer and orbital operator energy with hard cuts between angles, color-negative film rendition with fine grain, [stage-lighting color cast], heavy volumetric haze, real sweat sheen, 24fps 180° shutter, [XX] seconds.`
|
||||||
|
|
||||||
|
M5: `Camera Capture: wide-latitude cinema capture, vintage [XX]mm 2x anamorphic character at a wide aperture — oval bokeh, soft edge falloff — light diffusion bloom softening highlights, locked-off or extremely slow push-in only, color-negative film rendition with fine grain, palette grade [hex values], atmospheric haze, weathered material detail, 24fps 180° shutter, [XX] seconds. No humans, environment is the subject.`
|
||||||
|
|
||||||
|
## STACKING MODES (Multi-World Sequences)
|
||||||
|
|
||||||
|
Cuts between two worlds (M2-Void ↔ M1-Küche, M3 ↔ M4): write each shot's Camera Capture specs inline in the closing line — don't blend modes into one averaged grade; the cut IS the visual punch. Same-mode sequences: one continuous prompt, hard-cut triggers in Movement, one Camera Capture line with per-shot lens differences.
|
||||||
|
|
||||||
|
## LENS LENGTH GUIDE
|
||||||
|
|
||||||
|
32/35/40mm wide establishing, full-body, group · 50/55mm medium portrait, two-shot, waist-up · 75mm tight portrait, isolation, performance close-up · 85/100mm extreme close-up (eyes, lips, jewelry, fabric). Default: 55mm (M1/M3/M4), 50mm (M2); M5 eher 35→55mm.
|
||||||
|
|
||||||
|
## FRAME RATE NOTES
|
||||||
|
|
||||||
|
All modes default 24fps with 180° shutter. Slow-motion beats: `intercut 96fps high-speed slow-motion at [moment] holding 180° shutter.` — base stays 24fps.
|
||||||
|
|
||||||
|
## RUNTIME & PER-SHOT TIMING
|
||||||
|
|
||||||
|
Runtime stated in three places (title, Frame Map bei Sequenzen, Camera Capture) — all must match. Always ask, never default. Guidance: 4–8s one strong action · 8–12s action + reveal/hold · 12–15s 2–3 beats with hard cuts · complex sequences → separate prompts. Per-shot timing must sum to total.
|
||||||
|
|
||||||
|
## NEGATIVE → POSITIVE REWRITES
|
||||||
|
|
||||||
|
| Instinct (negative) | Lock (positive) |
|
||||||
|
|---|---|
|
||||||
|
| Don't change face | @image1 keeps the same face, hair, wardrobe, and silhouette throughout. |
|
||||||
|
| Don't switch positions | @image1 remains in the left third throughout; @image2 remains in the right third throughout. Neither crosses the center line. |
|
||||||
|
| Don't drift | Boots stay planted on the same ground marks across the full runtime. Only breath, eyes, hair, and fabric move subtly. |
|
||||||
|
| Don't change costume | Wardrobe identical across the runtime. |
|
||||||
|
| No extra people | The frame contains only @image1 and @image2 in their specified positions. No other figures enter or pass through. |
|
||||||
|
| No on-screen text | No on-screen text, no captions, no signage typography, no rendered text in the frame. |
|
||||||
|
| No camera chaos | Slow controlled handheld with natural operator breath, preserving @image1 in the left third and @image2 in the right third throughout. |
|
||||||
|
| No blur | Subjects remain sharply focused; controlled cinematic motion blur appears only on falling rain and distant background light sources. |
|
||||||
|
| Don't blink mid-action | Gaze stays locked on @image2 across the full runtime, eyes steady, no break in eye contact. |
|
||||||
|
| No mode switching | The shot runs as one continuous take with no cuts, no scene change, no time jump. |
|
||||||
|
|
||||||
|
Always prefer the positive form. Negative phrasing only in the explicit suppression lines for known failure modes.
|
||||||
|
|
||||||
|
## PRE-DELIVERY PASS (Silent QA)
|
||||||
|
|
||||||
|
Character gate asked and carried · every reference listed in check + bullet list + `@imageN`, order matching · canonical reference for every named subject (never substituted by the plate) · mode selected with rationale · Frame Map written · Subject Lock per character (wardrobe NOT re-described) · Cross-Frame Rules if 2+ · Movement with four layers and timestamps · Last Frame with suppression line · World Plate · Sound Bed diegetic · Capture Realism tuned (wet only if wet, skin only if humans, no gear overlap) · Camera Capture single line at bottom · lens chosen · runtime confirmed and matching · per-shot timing sums · no names/brands/tool names/context/meta · no music · English only · three-part delivery · all ten blocks in exact order · every reference tagged · negatives rewritten positive · word count in range.
|
||||||
|
|
||||||
|
**Repair pass:** too poetic → physical visual instructions · overloaded → split into sequence · drift risk → tighten Subject Lock (contact points, ground marks) · swap risk → tighten Cross-Frame Rules · wardrobe re-described → cut, trust the reference · double camera spec → collapse to one line · mode conflict → one dominant mode · action too complex → one dominant motion · Last Frame vague → specific closing composition · over word count → trim Subject Lock and Movement first, then Cross-Frame Rules.
|
||||||
|
|
||||||
|
## OPTIONAL HANDOFF — BANANA PRO DIRECTOR
|
||||||
|
|
||||||
|
If the user pairs a Seedance prompt with an existing Banana Pro plate: ask which cinema mode the plate used and lock the matching camera grammar — the two skills share the same five-mode framework; paired, still and video share visual DNA. Otherwise don't bring it up.
|
||||||
Reference in New Issue
Block a user