Real-media validation record
Cette page n’est pas encore traduite : voici la version anglaise.
What each workflow has been run on with real footage (not synthetic test patterns), what passed, and the
known limits. Sources are anonymised. All outputs stayed in a temp folder and none were committed: no real
media, transcripts, names or paths are in this repo. Synthetic regression tests live in tests/ and
workflows/*/tests/ (run python3 -m pytest tests -q for the current count). Per-workflow details:
workflows/<name>/PORT_NOTES.md.
Legend: Passed = checked by looking at stills and/or measuring; Limit = known gap or failure.
talkinghead
Section titled “talkinghead”Tested on
- The creator’s vertical HLG 口播 (iPhone, 1080x1920), first 65 s.
- The creator’s horizontal webcam recording (1080p, glasses), 60 s.
Passed
- HDR→SDR via avconvert; whisper word timestamps; cut_pass1 (65 → 51.5 s).
retouch_video.py(dailypreset, 6 workers) on a real face with thin metal glasses. Makeup flicker measured on a 180-frame 1080p clip: frame-to-frame lip / cheek luma change 1.01x / 0.94x the source’s, colour 1.00x, and no extra jump at chunk seams. Video makeup is on by default (natural, 0.3).- Retouch speed presets on a real vertical frame:
quality~0.64 s/frame,fast~0.21 s/frame (1 worker). Mean difference between them is 1.9 levels. - face_track hit 100 %. compose previews on 小红书 3:4 and 抖音 9:16, with hook, callout, 记笔记 panel, refined progress bar and pop word.
- Webcam reframe to 9:16, face mode: hit rate 1.0, p95 pan 0.05 crop-widths/s, 1.78x upscale recorded.
Checks on the frame:
- punch-in correctly capped to none (the face already filled 0.75 of the width);
- hook title shrunk above the glasses;
- PiP shrunk off the eyes;
- split and circle scenes OK.
- Webcam kept at 16:9 (
ORIENT=horizontal): YouTube landscape layout previews OK. - Split-screen B-roll on the webcam 抖音 preview. Captions had landed on the speaker’s mouth / chin; they now move to the B-roll pane above the seam (fixed after review).
Limit / found
-
Applying every strict_pass suggestion deleted real words, and the subtitles no longer matched the audio. Fixed by policy:
- only high-confidence fillers and repeats are pre-filled;
- the creator confirms the rest;
strict_pass.py verifyre-transcribes the cut and flags lost content words.
The verify step itself has not yet been re-run on the real clip.
-
The ffmpeg zscale HDR path has not been verified on real HLG (only avconvert has).
-
No full-length final render of real footage (previews and bodies only).
-
2026-10-05 demo round (a full real 口播 edit, 4:38 → 2:38, 3:4 + 9:16 masters + exports): re-applying strict after a drop mis-cut the body (passes now derive from immutable
segs.<stage>.jsonand refuse stale bodies); pass-1 edges pulled in neighbour syllables (now bounded by the neighbouring whisper words);cover.py --platformoverwrote the 3:4 cover; exports from a 9:16 master put 抖音 captions over burned panels (cues.json now carries keep-outs + keyword markup). Fixes regression-tested on synthetic media only (tests/test_talkinghead_fixes.py).
call-clips
Section titled “call-clips”Tested on
- A 3-person Zoom podcast, 60 s window.
Passed
render_trioon 小红书 9:16. Two non-creator tiles masked with a sticker, tracked at 100 % on both guests, andverify_coveragePASS 100 %.- All Zoom name labels blurred; checked unreadable before and after the crop.
- Auto-trim (classic profile) plus a 1.2x body gives 49.1 s, at −14.1 LUFS / −1.4 dBTP.
- Captions inside the caption box; panel and chips inside the safe box; 1080x1440 cover.
Limit / found
- Name chips were readable until
name_maskwas added. - 2-line captions climbed onto the host tile until the caption block was capped to the box height.
- faster-whisper backend not run on real media.
- 2026-10-05 demo round (a real 3-person podcast, 51 s 小红书 9:16 clip): the classic editor-cut snap ate the
previous word’s tail (now clamped at word ends), the hook→body dissolve faded the hook’s last word (hook ends now
follow the energy tail; fade shortened when tight), a 4-bullet panel sat under the chips with no WARN (auto-fit +
WARN), and the 9:16 trio left the lower frame empty with the silent creator largest (new stage layout: active
speaker large, masks follow the tile). Fixes regression-tested on synthetic media only
(
tests/test_call_clips_fixes.py); the stage layout has not been re-rendered on real media.
longform-to-short
Section titled “longform-to-short”Tested on
- A screen-share lecture with a participant avatar tile (no camera), 3-minute window.
- Rendered as vertical slices at 1080x1440 and 1080x1920.
Passed
- Title band used (no camera found). The avatar tile, the participant name and the browser bookmark bar never appear in frame.
- Code and doc text readable at ~2.1x, with the main code column and line numbers kept.
- −14.06 LUFS / −1.41 dBTP; label-length warnings fire.
Limit / found
- 72 px 2-line captions overflowed the 3:4 caption band; now capped to 60 px.
- Speaker framing on a real camera tile has not been verified (this recording had only an avatar).
- The first ~50 s transcribed as English noise: an ASR language-detection issue. Set
language.
photo-story
Section titled “photo-story”Tested on
- Museum / travel set: 10 iPhone HEIC photos and 2 HLG clips.
- Music-only mode with a CC-BY track, at 小红书 3:4.
Passed
- 128 BPM detected; 11 cuts on bars; section boundaries moved onto music sections.
beats.verifymax 0.49 frames. - HEIC rotation correct; HDR clip tone-mapped; title inside the caption box; chapter card on the section change.
- A kept clip measured −14.3 LUFS / −1.45 dBTP.
- On the synthetic demo (not real media): the default 3:4 and 9:16 stills are identical to the previous output. Platform stills for 3:4, 9:16, 16:9 and 小红书 horizontal checked with the safe and caption boxes drawn.
Limit / found
- An ambient track with no steady pulse (p90 270 ms) has no usable grid, so cuts follow the raw beats.
- Voice clone was run for real only with a synthetic reference voice, not the creator’s.
- No full-length render; 16:9 covers are still portrait-tuned.
lib (shared)
Section titled “lib (shared)”- 2026-10-05 demo round (six real projects): fixed in the lib - a MediaPipe VIDEO timestamp crash on face
re-acquire (2 of 8 retouch chunks), whisper captions invented over music-only clips (
asr.drop_hallucinations,has_speech), wav-then-AAC true peak -1.2 vs -1.5 dBTP, persona tags forced onto off-topic posts, a missing CJK serif role, same-aspect exports reported as letterbox, covers centre-cropped across aspects, re-burned captions over burned panels, and “no steady grid” on steady 128 BPM tracks with jittery beats. Each has a synthetic test intests/test_demo_round_lib.py.
Workflows validated on synthetic media only (no real-footage run yet)
Section titled “Workflows validated on synthetic media only (no real-footage run yet)”- vlog (calm + fun): synthetic drone-like clips and a drum track.
- Not run on real faces, real speech, HDR in the fun path, or mastered music.
- Fun compositing runs at ~10 output fps at 1080x1920.
- promo-recut: synthetic talk.
- Not run: a full HyperFrames render,
--verifywith real whisper, the face path of the vertical layout, or a retouched cover from real footage. The cover’s retouch code path runs, but on a face-less test image.
- Not run: a full HyperFrames render,
- explainer: 16:9 regression, byte-identical. A synthetic 3-line vertical short passes lint, and its caption boxes were pixel-checked. No full render; 3:4 is not snapshotted.
- cover, slides, polish, preproduction: synthetic inputs, pixel / command identical to the previous
version.
- Slides recording (Playwright) is unverified.
- The RVM matting engine was not verified on a real face.
- 2026-10-05 demo round: fun vlog speech gate (hallucinated ASR over template music rejected by transcript filter + loudness evidence; decisions in the report), cross-kind tag collisions, explicit/photo hook-finale-outro, A5 leak cap, map labels, face-aware scrapbook cover at exact size; photo-story loud serif fallback, POST tag options, per-shot minimum units in music mode, rotation-aware clip size. Synthetic tests: tests/test_vlog_photostory_fixes.py. Not re-run on the real demo footage.
- 2026-10-05 demo round: longform-to-short sample-exact render (PCM joins, one AAC encode; |A−V| < 1 frame over 30 items), midpoint caption words at splits, per-episode hooks, text-density zoom centres, merged vertical manifests, pad_in word guard, split layout filling the 3:4 / 9:16 frame with a readable-text zoom (synthetic share frames); explainer compositions/ mkdir, display_en comma / “a million” numbers, per-sentence TTS check with re-takes, suffixed per-platform cover names, vertical fill rule + layout_check.py. Synthetic tests: tests/test_lfs_explainer_fixes.py. Not re-run on the real demo footage / a full HyperFrames render.
- 2026-10-05 shared speech cleanup wiring: vlog fun speech shots (gentle by default, auto edits only, captions via
cleanup.remap_words), vlog calmcleanupsegments (analyze –ranges → apply before assembly), photo-story opt-incleanup=/CLEANUPonaudio="keep"clips, polish--cleanup pauses|gentle|standard|tight(off by default, never re-cuts a cleaned file). Synthetic tone-burst speech: tests/test_speech_cleanup_wiring.py. Not run on real footage / real whisper.