Real-media validation record
What each workflow has been run on with real footage (not synthetic test patterns), what passed, and the
known limits. Sources are anonymised. All outputs stayed in a temp folder and none were committed: no real
media, transcripts, names or paths are in this repo. Synthetic regression tests live in tests/ and
workflows/*/tests/ (run python3 -m pytest tests -q for the current count). Per-workflow details:
workflows/<name>/PORT_NOTES.md.
Legend: Passed = checked by looking at stills and/or measuring; Limit = known gap or failure.
talkinghead
Section titled “talkinghead”Tested on
- The creator’s vertical HLG 口播 (iPhone, 1080x1920), first 65 s.
- The creator’s horizontal webcam recording (1080p, glasses), 60 s.
Passed
- HDR→SDR via avconvert; whisper word timestamps; cut_pass1 (65 → 51.5 s).
retouch_video.py(dailypreset, 6 workers) on a real face with thin metal glasses. Makeup flicker measured on a 180-frame 1080p clip: frame-to-frame lip / cheek luma change 1.01x / 0.94x the source’s, colour 1.00x, and no extra jump at chunk seams. Video makeup is on by default (natural, 0.3).- Retouch speed presets on a real vertical frame:
quality~0.64 s/frame,fast~0.21 s/frame (1 worker). Mean difference between them is 1.9 levels. - face_track hit 100 %. compose previews on 小红书 3:4 and 抖音 9:16, with hook, callout, 记笔记 panel, refined progress bar and pop word.
- Webcam reframe to 9:16, face mode: hit rate 1.0, p95 pan 0.05 crop-widths/s, 1.78x upscale recorded.
Checks on the frame:
- punch-in correctly capped to none (the face already filled 0.75 of the width);
- hook title shrunk above the glasses;
- PiP shrunk off the eyes;
- split and circle scenes OK.
- Webcam kept at 16:9 (
ORIENT=horizontal): YouTube landscape layout previews OK. - Split-screen B-roll on the webcam 抖音 preview. Captions had landed on the speaker’s mouth / chin; they now move to the B-roll pane above the seam (fixed after review).
Limit / found
-
Applying every strict_pass suggestion deleted real words, and the subtitles no longer matched the audio. Fixed by policy:
- only high-confidence fillers and repeats are pre-filled;
- the creator confirms the rest;
strict_pass.py verifyre-transcribes the cut and flags lost content words.
The verify step itself has not yet been re-run on the real clip.
-
The ffmpeg zscale HDR path has not been verified on real HLG (only avconvert has).
-
No full-length final render of real footage (previews and bodies only).
-
2026-10-05 demo round (a full real 口播 edit, 4:38 → 2:38, 3:4 + 9:16 masters + exports): re-applying strict after a drop mis-cut the body (passes now derive from immutable
segs.<stage>.jsonand refuse stale bodies); pass-1 edges pulled in neighbour syllables (now bounded by the neighbouring whisper words);cover.py --platformoverwrote the 3:4 cover; exports from a 9:16 master put 抖音 captions over burned panels (cues.json now carries keep-outs + keyword markup). Fixes regression-tested on synthetic media only (tests/test_talkinghead_fixes.py).
call-clips
Section titled “call-clips”Tested on
- A 3-person Zoom podcast, 60 s window.
Passed
render_trioon 小红书 9:16. Two non-creator tiles masked with a sticker, tracked at 100 % on both guests, andverify_coveragePASS 100 %.- All Zoom name labels blurred; checked unreadable before and after the crop.
- Auto-trim (classic profile) plus a 1.2x body gives 49.1 s, at −14.1 LUFS / −1.4 dBTP.
- Captions inside the caption box; panel and chips inside the safe box; 1080x1440 cover.
Limit / found
- Name chips were readable until
name_maskwas added. - 2-line captions climbed onto the host tile until the caption block was capped to the box height.
- faster-whisper backend not run on real media.
- 2026-10-05 demo round (a real 3-person podcast, 51 s 小红书 9:16 clip): the classic editor-cut snap ate the
previous word’s tail (now clamped at word ends), the hook→body dissolve faded the hook’s last word (hook ends now
follow the energy tail; fade shortened when tight), a 4-bullet panel sat under the chips with no WARN (auto-fit +
WARN), and the 9:16 trio left the lower frame empty with the silent creator largest (new stage layout: active
speaker large, masks follow the tile). Fixes regression-tested on synthetic media only
(
tests/test_call_clips_fixes.py); the stage layout has not been re-rendered on real media.
longform-to-short
Section titled “longform-to-short”Tested on
- A screen-share lecture with a participant avatar tile (no camera), 3-minute window.
- Rendered as vertical slices at 1080x1440 and 1080x1920.
Passed
- Title band used (no camera found). The avatar tile, the participant name and the browser bookmark bar never appear in frame.
- Code and doc text readable at ~2.1x, with the main code column and line numbers kept.
- −14.06 LUFS / −1.41 dBTP; label-length warnings fire.
Limit / found
- 72 px 2-line captions overflowed the 3:4 caption band; now capped to 60 px.
- Speaker framing on a real camera tile has not been verified (this recording had only an avatar).
- The first ~50 s transcribed as English noise: an ASR language-detection issue. Set
language.
photo-story
Section titled “photo-story”Tested on
- Museum / travel set: 10 iPhone HEIC photos and 2 HLG clips.
- Music-only mode with a CC-BY track, at 小红书 3:4.
Passed
- 128 BPM detected; 11 cuts on bars; section boundaries moved onto music sections.
beats.verifymax 0.49 frames. - HEIC rotation correct; HDR clip tone-mapped; title inside the caption box; chapter card on the section change.
- A kept clip measured −14.3 LUFS / −1.45 dBTP.
- On the synthetic demo (not real media): the default 3:4 and 9:16 stills are identical to the previous output. Platform stills for 3:4, 9:16, 16:9 and 小红书 horizontal checked with the safe and caption boxes drawn.
Limit / found
- An ambient track with no steady pulse (p90 270 ms) has no usable grid, so cuts follow the raw beats.
- Voice clone was run for real only with a synthetic reference voice, not the creator’s.
- No full-length render; 16:9 covers are still portrait-tuned.
lib (shared)
Section titled “lib (shared)”- 2026-10-05 demo round (six real projects): fixed in the lib - a MediaPipe VIDEO timestamp crash on face
re-acquire (2 of 8 retouch chunks), whisper captions invented over music-only clips (
asr.drop_hallucinations,has_speech), wav-then-AAC true peak -1.2 vs -1.5 dBTP, persona tags forced onto off-topic posts, a missing CJK serif role, same-aspect exports reported as letterbox, covers centre-cropped across aspects, re-burned captions over burned panels, and “no steady grid” on steady 128 BPM tracks with jittery beats. Each has a synthetic test intests/test_demo_round_lib.py.
Workflows validated on synthetic media only (no real-footage run yet)
Section titled “Workflows validated on synthetic media only (no real-footage run yet)”- vlog (calm + fun): synthetic drone-like clips and a drum track.
- Not run on real faces, real speech, HDR in the fun path, or mastered music.
- Fun compositing runs at ~10 output fps at 1080x1920.
- promo-recut: synthetic talk.
- Not run: a full HyperFrames render,
--verifywith real whisper, the face path of the vertical layout, or a retouched cover from real footage. The cover’s retouch code path runs, but on a face-less test image.
- Not run: a full HyperFrames render,
- explainer: 16:9 regression, byte-identical. A synthetic 3-line vertical short passes lint, and its caption boxes were pixel-checked. No full render; 3:4 is not snapshotted.
- cover, slides, polish, preproduction: synthetic inputs, pixel / command identical to the previous
version.
- Slides recording (Playwright) is unverified.
- The RVM matting engine was not verified on a real face.
- 2026-10-05 demo round: fun vlog speech gate (hallucinated ASR over template music rejected by transcript filter + loudness evidence; decisions in the report), cross-kind tag collisions, explicit/photo hook-finale-outro, A5 leak cap, map labels, face-aware scrapbook cover at exact size; photo-story loud serif fallback, POST tag options, per-shot minimum units in music mode, rotation-aware clip size. Synthetic tests: tests/test_vlog_photostory_fixes.py. Not re-run on the real demo footage.
- 2026-10-05 demo round: longform-to-short sample-exact render (PCM joins, one AAC encode; |A−V| < 1 frame over 30 items), midpoint caption words at splits, per-episode hooks, text-density zoom centres, merged vertical manifests, pad_in word guard, split layout filling the 3:4 / 9:16 frame with a readable-text zoom (synthetic share frames); explainer compositions/ mkdir, display_en comma / “a million” numbers, per-sentence TTS check with re-takes, suffixed per-platform cover names, vertical fill rule + layout_check.py. Synthetic tests: tests/test_lfs_explainer_fixes.py. Not re-run on the real demo footage / a full HyperFrames render.
- 2026-10-05 shared speech cleanup wiring: vlog fun speech shots (gentle by default, auto edits only, captions via
cleanup.remap_words), vlog calmcleanupsegments (analyze –ranges → apply before assembly), photo-story opt-incleanup=/CLEANUPonaudio="keep"clips, polish--cleanup pauses|gentle|standard|tight(off by default, never re-cuts a cleaned file). Synthetic tone-burst speech: tests/test_speech_cleanup_wiring.py. Not run on real footage / real whisper.