Guillemard Terrace — AI listing reel, end-to-end build log

A 12.4-second vertical listing reel generated entirely from still photos, one reference person, and a recorded voice line. This page documents every input, every prompt, and every post-processing step that produced the final video. The interactive shot plan is at shotplan.html.

The finished video

Final cut — 10.79s, 480×854 (9:16), narration + music + captions burned in.

How it was made, step by step

1
Collect the real inputs. Three listing photos of the house (front facade, side garden + deck, dining room under the stair void), one reference photo of the presenter Ernest, a screenshot showing his intended outfit, and a voice recording of the line "This stunning 3-story home is yours for just 3.67 million." Nothing else — no footage was shot for this video.
2
Lock the facts. Before any prompt was written, the system recorded what each photo actually shows (folding doors, grey deck, two curved balconies, red ixora pot, floating staircase, etc.) and what camera moves each photo can support. This contract is what keeps the AI truthful — prompts may only move through space a photo actually captured.
3
Measure the narration. The voice line was transcribed word-by-word with exact timings (speech runs 0.00–3.96s inside a 4.31s clip). Those timings drive everything downstream: the acting beat, the camera snap, the captions, and the edit points.
4
Build Ernest's character sheet. A three-panel identity sheet was generated through the avatar-realism identity pipeline: left panel headless front wardrobe, middle panel full back view, right panel three-quarter face close-up — one outfit and one identity locked across panels. It went through the pipeline's compile → preflight → provider render → independent visual review → finalization gates (state: locally_verified). This sheet is what the video model uses as "@Image2 — Ernest".
5
Write the A-roll prompt. The presenter prompt describes one continuous phone-camera take: Ernest stands on the deck in front of the folding doors, mic held at his chin in his left hand, and on the word "million" the camera snap-zooms out to a 0.5× ultra-wide lens, revealing the whole three-storey facade. (The prompt text below is the later revision that adds a 2-second held wide frame — the final cut uses the first render, made before that revision.) The prompt pins his face, wardrobe, lip-sync to the recorded voice, and forbids the camera from inventing anything outside the facade photo.
6
Write the B-roll prompt. A six-second pack of three 2-second shots across two photos: (i) side garden — slow push along the stepping stones toward the covered deck; (ii) side garden — slow orbit around the red ixora pot; (iii) dining room — push toward the under-stair water garden. Hard cuts between shots; each shot stays inside what its source photo shows.
7
Render two clips. Both clips were generated by Seedance 2.5 (omni-reference mode) through the Higgsfield API — 9:16 vertical, 480p, no generated audio. The A-roll got the facade photo, the character sheet, a black-screen video carrying the voice recording, and first/last frame plates. The B-roll got the two room photos.
8
Stitch. The first A-roll render plays out to its natural end (5.04s — no early cut), then a short 0.3s dissolve hands off to the B-roll, which plays all 6 seconds — a 10.79s picture track.
9
Attach the music. The instrumental "Free This Feeling" rides as a quiet bed under the speech, then swells +10 dB the moment Ernest stops talking and runs to the end. A cash-register "kaching" lands just after the word million, on the zoom reveal. The mix follows the project's committed audio recipe (loudness-normalized to −16 LUFS) — the earlier 10s master had a bug (amix duration=first truncated it to the narration), which was found and fixed here.
10
Burn the captions. Phrase captions timed to the measured word timings — five short chunks, with the price line in the accent color — rendered into the video itself (always visible, no player needed).

Timeline

video
A-ROLL · 0–5.04s (plays out whole)
0.3s dissolve → B-ROLL · 4.74–10.79s (3 shots)
voice
narration 0–3.96s
music
quiet bed (ducked under speech)
+10 dB swell 4.31–10.79s
sfx / captions
kaching @4.0s
captions 0–3.96s

Per-photo screen time is marked lane-by-lane in the shot plan: facade 0–5.04s · side-garden 4.74–8.74s · dining-void 8.74–10.79s.

Every input, and what it fed

stills/facade.jpg
Listing photo → @Image1 of the A-roll: the location the camera zooms out across.
stills/side-garden-deck.jpg
Listing photo → @Image1 of the B-roll: shots 1–2 (push to deck, orbit ixora).
stills/dining-void.jpg
Listing photo → @Image2 of the B-roll: shot 3 (push to water garden).
agent/ernest-ref.jpg
Ernest's face photo → identity input for the character sheet.
outfit screenshot
Supplied wardrobe/build reference → outfit input for the character sheet.
agent/ernest-identity-sheet.png
Verified 3-panel character sheet → @Image2 of the A-roll (who Ernest is).
plates/a-start-composite.png
Opening-frame plate → Seedance start frame.
plates/a-end-composite.png
Closing-frame plate (wide facade) → Seedance end frame.
plates/b-garden-9x16.jpg
9:16 crop of the side-garden photo → B-roll framing reference.

Audio inputs (not pictured): run/narration/hook.wav (the voice line, 4.31s, sha-pinned), muxed into carrier-hook-hold.mp4 — a 7s black-screen clip so the video model receives the voice as a "video reference" whose audio is law. Music: free-this-feeling-instrumental.mp3 (seek 50s). SFX: price-kaching.mp3.

The two renders

A-roll — job cf5238e2, 5s. Ernest delivers the line; camera snap-zooms on "million" to reveal the full facade. (A later 7s variant with a 2s held wide frame — job 13a82b58 — was rendered but the final cut uses this first render played to its natural end.)
B-roll — job cf3e6ad1, 6s. Three shots: deck push → ixora orbit → water-garden push.

Properties

PropertyValue
Modelseedance_2_5 (ByteDance), omni_reference mode, via Higgsfield CLI
Format9:16 vertical · 480×854 · 24 fps · h264 + aac
A-roll5s container, plays out whole · refs: facade photo, identity sheet, speech-carrier video, start/end plates · job cf5238e2
B-roll6s container, used whole · refs: 2 room photos · job cf3e6ad1
Narration12 words, speech 0.00–3.96s in a 4.31s stem; zero offset on the master
Musicfree-this-feeling-instrumental · bed −26 dB under speech · +10 dB lift 4.31–10.79s · 1.5s fade-out
SFXprice-kaching at 4.0s (lands just after "million" ends at 3.96s)
Captions5 phrase chunks from the committed word alignment, Avenir Next bold, burned in
Mixcommitted mix_recipe semantics · loudnorm I=−16, TP=−1.5 · run/audio/master-11s-mixed.wav
Truth locksevery shot stays inside a source photo; no invented rooms, objects or camera moves; presenter identity pinned to the verified sheet

Word timings (drives acting, edit and captions)

wordstartend
This0.00s0.32s
stunning0.32s0.62s
30.62s0.94s
-story0.94s1.18s
home1.18s1.54s
is1.54s1.80s
yours1.80s2.10s
for2.10s2.40s
just2.40s2.66s
32.66s3.02s
.673.02s3.62s
million.3.62s3.96s

A-roll prompt (exact text sent to the model)

【GENERATION GOAL】
A wide establishing listing presentation: Ernest presents this home's front elevation to a phone camera; on the word «million» the camera snaps to its 0.5 ultra-wide lens, revealing the facade at its full three-storey height, then holds that wide frame for the final two seconds. Generate the picture; the speech already exists in @Video1. This is a generation, not an edit of @Video1.

【REFERENCE ASSET ROLES】
@Image1 — the FACADE: the photographed three-storey front elevation — folding doors, grey deck, two curved balconies, curved roof, car porch, parked cars, frangipani tree, overcast daylight. It supplies geography, architecture, materials and light. The person visible inside through the right-hand door glass is not part of this shot.
@Image2 — ERNEST: his character sheet — that face and short black hair, with the build and wardrobe shown in its body panels: an oversized black t-shirt with a small gold EVISU emblem, loose black wide-leg trousers, and black shoes. The small black mic is an authored prop held in his left hand, not part of the sheet. The sheet's background and panel layout are not part of this shot.
@Video1 — the SPEECH CARRIER: a narration file whose audio is the master. Only its audio enters this generation. Nothing from its picture, identity, wardrobe or scene enters this shot.

【SUBJECTS】
Ernest maps to @Image2 — that face, short black hair, oversized black t-shirt, black wide-leg trousers, black shoes, small handheld mic in the left hand. One Ernest, existing once.
The scene is @Image1 and nothing else: black-framed glass folding doors at ground centre; grey composite deck in front; car porch under a curved metal roof with parked cars at frame-left; frangipani over a low wall at frame-right; two stacked curved balconies under a curved roof; pebble-and-lawn strip before the deck. Every object in frame appears in @Image1 or on Ernest's reference. One facade.

ENVIRONMENTAL MOTION CONTRACT — No added environmental motion. Preserve the scene exactly as shown in @Image1; do not invent a fan, airflow, moving object, changing shadow, or lighting change. Ernest and the presenter's own cast shadow are exempt and move naturally.

【THE SPEECH — @Video1 IS THE SOUND LAW】

CARRIER: @Video1 is a narration file. Its picture never appears — every frame of the output shows the facade and Ernest at full brightness, first frame to last, opening and closing on the picture with no fade.

MASTER AUDIO: the audio of @Video1 is used whole and untouched, start to finish — same voice, same timing, same pauses, same words. Nothing is re-read, re-voiced, cloned, regenerated, re-timed, stretched or compressed. Picture and audio start together at 0:00 and end together.

WORDS OF THIS BLOCK (as heard in @Video1 — the FILE is the truth): "This stunning 3-story home is yours for just 3.67 million."

MEASURED FLOW (from the file): the file opens on «This» at 0:00 — the mouth is live on the first word; the last word «million» ends at ≈3.96s; the remaining ≈0.35s of the 4.31s stem is quiet.

MOUTH OWNERSHIP: every audible word belongs to Ernest, and that is the only mouth in frame. The presenter speaks these words and nothing else — no ad-libs, no "uhm", no second voice, no echo.

MIC LAW: the grille rides at the presenter's collarbone, below the lips, and never rises across the mouth. In every frame where the presenter speaks, the mouth is unobstructed.

TAIL: when the last word ends, the presenter's mouth closes and the face stays alive — the eyes hold the lens, the blink cadence continues — and he holds his post, planted and still, through the wide reveal to the final frame. The picture outlasts the spoken stem by about two seconds; it does not cut with the last word.

FORMAT — ONE CONTINUOUS TAKE. The speech runs the full length of @Video1, and the picture continues about two seconds past the stem's end, holding the wide frame in silence. Real time.

LOCATION MAP — Compass: frame left=curved car-porch roof with parked cars; frame right=frangipani over the low boundary wall; upstage behind=the facade elevation. STAGE is the photographed grey composite deck directly in front of the centre folding doors. All coverage shoots from this position and stays on this side of the lens axis, so frame-left stays frame-left for every shot in this scene.

FRAMING — Full-length wide in 9:16 vertical. Ernest stands at medium scale in the opening frame — crown below the balcony band above the doors, feet on the deck boards, about 60-65% of frame height — centred on the doors. The lower portion reads the source-visible deck boards and pebble-and-lawn strip below and around the presenter's feet; the upper frame keeps generous source-visible headroom under the balconies and curved roof. The visible scene is limited to this source inventory: folding doors; grey deck; car porch with parked cars; frangipani and low wall; two curved balconies; curved roof. Keep only edges, structure and floor that @Image1 actually shows; never invent an object, surface, wall or depth to satisfy the composition. @Image1 remains the authority for geometry, materials, light and atmosphere.

OFF-SCREEN LOCK — Ernest is the only person outside in frame. The person visible inside through the right-hand door glass stays inside and never acknowledges the camera. The operator and the phone stay outside it: no limb, no shadow, no reflection in any glazing. The deck carries only what @Image1 shows. The camera never pans, dollies, or reframes into a region @Image1 does not show: no invented second window, extra wall, or architecture outside the photograph. A move that would open a blindspot is not a legal move; the lens stays inside the photographed scene.

HEIGHT RULER — In the opening frame Ernest's crown sits below the balcony band above the doors, feet on the deck boards, about 60-65% of frame height. In the closing frame he reads small — about 20% — standing at the same spot.

FIRST FRAME AND BLOCKING — Ernest starts at the verified placement point on the grey composite deck directly in front of the centre folding doors; both feet contact only the deck boards. Do not place either foot or any body contact off the photographed deck plane. Photographic orientation anchors: frame left: car-porch roof and parked cars; frame right: frangipani and low wall; upstage behind: folding doors and balconies. The carrier alone determines mouth state at frame zero; the mouth is at rest on frame one and opens on «This».

MOVEMENT CONTRACT — Natural body movement follows the speech timing and stays within the photographed deck. Remain at the verified placement point on the grey composite deck in front of the folding doors; no walking or travel. VISIBLE ACTION — the small handheld microphone stays gripped in the left hand, held at the chin through the whole line; the right hand gestures through the spoken phrases, and on the word «million» the right hand shoots out into an open palm — thrown wide to express how huge the home is — while the left hand keeps the mic at the chin and he keeps speaking into it. This exact authored action is the sole authority for the free hand and gaze; do not add illustrative gestures. The mic is never clipped on, transferred, or replaced by a hidden lavalier. LIP-SYNC LAW owns the head. Blinks, breath, micro-saccades and subtle weight transfer remain naturally alive through the final words, the quiet tail, and the two-second held wide frame.

CAMERA — Handheld phone at chest height, roughly 1.5 metres off the photographed ground, real time, height constant and dead level, with no tilt. Slight shake from small operator corrections, the horizon steady. The camera never pans or reframes into a region @Image1 does not show, and its only move is scripted: on the word «million» the camera executes ONE snap zoom-out to the phone's 0.5 ultra-wide lens — a fast step wider, not a slow dolly or pull — instantly widening until the frame holds the FACADE at its full three-storey height: porch and cars frame-left, both balconies and the curved roof top to bottom, the frangipani frame-right — revealing only what @Image1 already contains. After the snap the camera holds that wide composition locked for the remaining roughly two seconds of the take — settled framing with only the natural slight shake of a handheld phone; no re-aim, no drift, no second move. Natural phone rendering, straight lines straight at the frame edges. Reflections off the glass only where the scene's own light justifies them.

PHYSICS — The presenter's hair lags a beat behind each head and shoulder turn and settles. The t-shirt fabric moves with the shoulders. Weight transfers through the sneakers into the deck boards. The mic is wireless, no cable. Skin keeps its own texture and pores, and the picture keeps the references' own softness, contrast and colour.

LIGHTING — Match the photo's overcast daylight exactly: soft even light, grey sky, one soft shadow direction. No relight, no added sun, no beauty key, no rim light. Exposure holds on his face through the zoom reveal.

AUDIO — The voice is @Video1's, whole and untouched. Under it, quiet outdoor room tone — faint foliage, distant street. No BGM, no subtitles.

B-roll prompt (exact text sent to the model)

REFERENCES AND DURATION
Inputs: @Image1 (room listing photo; no depth map supplied); @Image2 (dining room listing photo; no depth map supplied) — 2 photographed viewpoints.
Duration: 6.0 seconds total, 3 shots, 2 hard cuts. 9:16 vertical.

SCENE CONTEXT
Daylight multi-room listing walkthrough covering room (@Image1) and dining room (@Image2); hard cuts between photographed rooms, no invented corridors.

ACTIVE REFERENCES
@Image1: source of truth for room geometry, furniture, materials, colors, views, and lighting.
No depth map is supplied for @Image1; infer no unseen geometry and stay inside the source photo's visible space.
@Image2: source of truth for dining room geometry, furniture, materials, colors, views, and lighting.
No depth map is supplied for @Image2; infer no unseen geometry and stay inside the source photo's visible space.

ROOM MAP
@Image1 ROOM:
- SCREEN-LEFT (Mid, Mid Depth): grass strip with palm trunks; laundry visible at the far end.
- SCREEN-RIGHT (Mid-to-Far Depth): grey deck under a grey awning with a tall black bar table and two stools, black-framed folding doors behind the deck, large red-brown pot of red-flowering ixora at the deck edge.
- CENTER (Foreground, Bright Depth): large grey stepping stones set in white pebbles leading to the deck, grey corrugated awning spanning the top of frame.
The photographed extent ends at the @Image1 screen-left boundary; nothing exists beyond it. The photographed extent ends at the @Image1 screen-right boundary; nothing exists beyond it. The photographed extent ends at the @Image1 screen-top boundary; nothing exists beyond it. The photographed extent ends at the @Image1 screen-bottom boundary; nothing exists beyond it.
@Image2 DINING ROOM:
- SCREEN-LEFT (Mid, Mid Depth): tall yellow-orange vase with bare branches by the left wall.
- SCREEN-RIGHT (Mid, Mid Depth): tall black-framed glass wall along the right side.
- FAR CENTER (Background, Dark Depth): open-tread timber staircase rising to a landing under roof glazing.
- CENTER (Foreground-to-Mid, Bright-to-Mid Depth): magenta-clothed dining table with dark timber floral-seat chairs, dark water feature with two white wrapped statuettes under the floating staircase.
The photographed extent ends at the @Image2 screen-left boundary; nothing exists beyond it. The photographed extent ends at the @Image2 screen-right boundary; nothing exists beyond it. The photographed extent ends at the @Image2 screen-top boundary; nothing exists beyond it. The photographed extent ends at the @Image2 screen-bottom boundary; nothing exists beyond it.
Reflective surfaces show reflections only; screens stay off and dark.

SOURCE FRAMING
Each interval is a full-height 9:16 crop of its source photo; discard width, never squeeze. Top/bottom are source frame boundaries; interval edges are named below. Keep circles, lines, and scale true.

POV AND VISIBILITY LOCK
Start at each source photo's entry viewpoint at standing eye height; move only through visible space in that photo. Keep behind-camera space and space beyond source edges out of view. No new doors, windows, walls, or rooms; no ceiling above or floor below the captured extent. Never interpolate through unphotographed space between rooms.

FORMAT MODE
CONTROLLED MULTI-SHOT SEQUENCE. New under-one-third-overlap crop at each cut; landmark per interval. HARD CUT only. No fades or dissolves.

CAMERA PATH
0s-2s: Fresh center full-height crop of @Image1; top-left vertical edge cuts the stepping stones in pebbles; farthest frame-right cuts the red ixora pot; screen-top/bottom are @Image1 source frame boundaries. Wide side-garden view along the stepping stones. Start: stepping stones in pebbles lower screen-left, grey awning underside upper screen-left, covered deck with bar table screen-center. Already moving in a slow gimbal walk toward the covered deck with bar table; keeping the covered deck with bar table in frame; do not fly to a new subject; move at one even constant speed from first frame to last, with no acceleration, deceleration, ease-in, ease-out, or speed ramp. End: covered deck with bar table remains at screen-right with surrounding room still visible.

At the 2-second mark: HARD CUT to @Image1; do not interpolate viewpoint.

2s-4s: Fresh right full-height crop of @Image1; top-left vertical edge cuts the folding glass doors; farthest frame-right is the @Image1 screen-right boundary; screen-top/bottom are @Image1 source frame boundaries. Standing at the deck edge beside the red ixora pot. Start: folding glass doors screen-left, red ixora pot screen-right. Slow orbit right around the red ixora pot; the camera travels a level constant-radius path around the red ixora pot while the lens stays trained on the red ixora pot; keeping the red ixora pot in frame; do not fly to a new subject; do not close in; move at one even constant speed from first frame to last, with no acceleration, deceleration, ease-in, ease-out, or speed ramp. End: red ixora pot remains at screen-right with surrounding room still visible.

At the 4-second mark: HARD CUT to @Image2; do not interpolate viewpoint.

4s-6s: Fresh center full-height crop of @Image2; top-left vertical edge cuts the dining table and chairs; farthest frame-right cuts the full-height glass wall; screen-top/bottom are @Image2 source frame boundaries. Dining-table foreground looking along the pebble strip. Start: dining table and chairs lower screen-left, under-stair water garden lower screen-left, floating staircase screen-right. Slow gimbal walk toward the under-stair water garden; keeping the under-stair water garden in frame; do not fly to a new subject; stop before its near face; move at one even constant speed from first frame to last, with no acceleration, deceleration, ease-in, ease-out, or speed ramp. End: under-stair water garden remains at lower screen-center with surrounding room still visible.

OPTICS
0s-2s: 84° wide rectilinear, deep focus; 2s-4s: 84° wide rectilinear, deep focus; 4s-6s: 84° wide rectilinear, deep focus. Level camera, straight verticals; no fisheye, dutch angle, or lens drift.

MOTION
Gimbal-stabilized and level, with no shake, sway, or footstep bounce. Each shot's camera move holds one even constant speed from first frame to last, bounded by the photographed space. No acceleration or deceleration. No ease-in. No ease-out. No speed ramp. The camera holds its stated move; photographed architecture stays true and stable. No zoom, optical or digital. Indoor: framing height and horizon stay constant for the whole shot.

LIGHTING
Match each source photo's daylight, shadows, color temperature, fixture states, curtains, and window state. Exposure stays constant; no relighting, grade, flare, or pumping.

AUDIO
No audio, no BGM, no subtitles.

FIDELITY LOCKS
Furniture, materials, colors, and floor pattern stay fixed to the active source photo.
Screens stay off with a dark blank screen.
Window views and reflections stay consistent.
No added furniture, décor, plants, rugs, or staging.
No added people. Do not invent pets or animals.
No on-screen text, logos, or watermarks.
Room proportions and scale stay fixed.

WORLD MOTION
If something already visible in the active source photo would move in real life, let it move at a natural pace. Do not freeze it. Do not add movers that are not already in the photo.
Cable cars or gondolas on a visible cable travel along that cable.
Vehicles on a visible road keep driving.
Boats or ships on visible water keep moving with the water.
Water has a light natural ripple where water is shown.
A fan or any other spinning object already in the photo keeps spinning.
A real pet already in the photo (dog, cat, bird) may move like a living animal. A stuffed toy, plush, statue, or printed animal stays still.
Do not invent cable cars, traffic, boats, people, or animals.

Built by the 1up video orchestrator + avatar-realism identity pipeline + housewalk B-roll conventions. Generated 2026-09-12 (SGT). Interactive shot plan: shotplan.html