Skip to content

Real-media validation record

What each workflow has been run on with real footage (not synthetic test patterns), what passed, and the known limits. Sources are anonymised. All outputs stayed in a temp folder and none were committed: no real media, transcripts, names or paths are in this repo. Synthetic regression tests live in tests/ and workflows/*/tests/ (run python3 -m pytest tests -q for the current count). Per-workflow details: workflows/<name>/PORT_NOTES.md.

Legend: Passed = checked by looking at stills and/or measuring; Limit = known gap or failure.

Tested on

  • The creator’s vertical HLG 口播 (iPhone, 1080x1920), first 65 s.
  • The creator’s horizontal webcam recording (1080p, glasses), 60 s.

Passed

  • HDR→SDR via avconvert; whisper word timestamps; cut_pass1 (65 → 51.5 s).
  • retouch_video.py (daily preset, 6 workers) on a real face with thin metal glasses. Makeup flicker measured on a 180-frame 1080p clip: frame-to-frame lip / cheek luma change 1.01x / 0.94x the source’s, colour 1.00x, and no extra jump at chunk seams. Video makeup is on by default (natural, 0.3).
  • Retouch speed presets on a real vertical frame: quality ~0.64 s/frame, fast ~0.21 s/frame (1 worker). Mean difference between them is 1.9 levels.
  • face_track hit 100 %. compose previews on 小红书 3:4 and 抖音 9:16, with hook, callout, 记笔记 panel, refined progress bar and pop word.
  • Webcam reframe to 9:16, face mode: hit rate 1.0, p95 pan 0.05 crop-widths/s, 1.78x upscale recorded. Checks on the frame:
    • punch-in correctly capped to none (the face already filled 0.75 of the width);
    • hook title shrunk above the glasses;
    • PiP shrunk off the eyes;
    • split and circle scenes OK.
  • Webcam kept at 16:9 (ORIENT=horizontal): YouTube landscape layout previews OK.
  • Split-screen B-roll on the webcam 抖音 preview. Captions had landed on the speaker’s mouth / chin; they now move to the B-roll pane above the seam (fixed after review).

Limit / found

  • Applying every strict_pass suggestion deleted real words, and the subtitles no longer matched the audio. Fixed by policy:

    • only high-confidence fillers and repeats are pre-filled;
    • the creator confirms the rest;
    • strict_pass.py verify re-transcribes the cut and flags lost content words.

    The verify step itself has not yet been re-run on the real clip.

  • The ffmpeg zscale HDR path has not been verified on real HLG (only avconvert has).

  • No full-length final render of real footage (previews and bodies only).

  • 2026-10-05 demo round (a full real 口播 edit, 4:38 → 2:38, 3:4 + 9:16 masters + exports): re-applying strict after a drop mis-cut the body (passes now derive from immutable segs.<stage>.json and refuse stale bodies); pass-1 edges pulled in neighbour syllables (now bounded by the neighbouring whisper words); cover.py --platform overwrote the 3:4 cover; exports from a 9:16 master put 抖音 captions over burned panels (cues.json now carries keep-outs + keyword markup). Fixes regression-tested on synthetic media only (tests/test_talkinghead_fixes.py).

Tested on

  • A 3-person Zoom podcast, 60 s window.

Passed

  • render_trio on 小红书 9:16. Two non-creator tiles masked with a sticker, tracked at 100 % on both guests, and verify_coverage PASS 100 %.
  • All Zoom name labels blurred; checked unreadable before and after the crop.
  • Auto-trim (classic profile) plus a 1.2x body gives 49.1 s, at −14.1 LUFS / −1.4 dBTP.
  • Captions inside the caption box; panel and chips inside the safe box; 1080x1440 cover.

Limit / found

  • Name chips were readable until name_mask was added.
  • 2-line captions climbed onto the host tile until the caption block was capped to the box height.
  • faster-whisper backend not run on real media.
  • 2026-10-05 demo round (a real 3-person podcast, 51 s 小红书 9:16 clip): the classic editor-cut snap ate the previous word’s tail (now clamped at word ends), the hook→body dissolve faded the hook’s last word (hook ends now follow the energy tail; fade shortened when tight), a 4-bullet panel sat under the chips with no WARN (auto-fit + WARN), and the 9:16 trio left the lower frame empty with the silent creator largest (new stage layout: active speaker large, masks follow the tile). Fixes regression-tested on synthetic media only (tests/test_call_clips_fixes.py); the stage layout has not been re-rendered on real media.

Tested on

  • A screen-share lecture with a participant avatar tile (no camera), 3-minute window.
  • Rendered as vertical slices at 1080x1440 and 1080x1920.

Passed

  • Title band used (no camera found). The avatar tile, the participant name and the browser bookmark bar never appear in frame.
  • Code and doc text readable at ~2.1x, with the main code column and line numbers kept.
  • −14.06 LUFS / −1.41 dBTP; label-length warnings fire.

Limit / found

  • 72 px 2-line captions overflowed the 3:4 caption band; now capped to 60 px.
  • Speaker framing on a real camera tile has not been verified (this recording had only an avatar).
  • The first ~50 s transcribed as English noise: an ASR language-detection issue. Set language.

Tested on

  • Museum / travel set: 10 iPhone HEIC photos and 2 HLG clips.
  • Music-only mode with a CC-BY track, at 小红书 3:4.

Passed

  • 128 BPM detected; 11 cuts on bars; section boundaries moved onto music sections. beats.verify max 0.49 frames.
  • HEIC rotation correct; HDR clip tone-mapped; title inside the caption box; chapter card on the section change.
  • A kept clip measured −14.3 LUFS / −1.45 dBTP.
  • On the synthetic demo (not real media): the default 3:4 and 9:16 stills are identical to the previous output. Platform stills for 3:4, 9:16, 16:9 and 小红书 horizontal checked with the safe and caption boxes drawn.

Limit / found

  • An ambient track with no steady pulse (p90 270 ms) has no usable grid, so cuts follow the raw beats.
  • Voice clone was run for real only with a synthetic reference voice, not the creator’s.
  • No full-length render; 16:9 covers are still portrait-tuned.
  • 2026-10-05 demo round (six real projects): fixed in the lib - a MediaPipe VIDEO timestamp crash on face re-acquire (2 of 8 retouch chunks), whisper captions invented over music-only clips (asr.drop_hallucinations, has_speech), wav-then-AAC true peak -1.2 vs -1.5 dBTP, persona tags forced onto off-topic posts, a missing CJK serif role, same-aspect exports reported as letterbox, covers centre-cropped across aspects, re-burned captions over burned panels, and “no steady grid” on steady 128 BPM tracks with jittery beats. Each has a synthetic test in tests/test_demo_round_lib.py.

Workflows validated on synthetic media only (no real-footage run yet)

Section titled “Workflows validated on synthetic media only (no real-footage run yet)”
  • vlog (calm + fun): synthetic drone-like clips and a drum track.
    • Not run on real faces, real speech, HDR in the fun path, or mastered music.
    • Fun compositing runs at ~10 output fps at 1080x1920.
  • promo-recut: synthetic talk.
    • Not run: a full HyperFrames render, --verify with real whisper, the face path of the vertical layout, or a retouched cover from real footage. The cover’s retouch code path runs, but on a face-less test image.
  • explainer: 16:9 regression, byte-identical. A synthetic 3-line vertical short passes lint, and its caption boxes were pixel-checked. No full render; 3:4 is not snapshotted.
  • cover, slides, polish, preproduction: synthetic inputs, pixel / command identical to the previous version.
    • Slides recording (Playwright) is unverified.
    • The RVM matting engine was not verified on a real face.
  • 2026-10-05 demo round: fun vlog speech gate (hallucinated ASR over template music rejected by transcript filter + loudness evidence; decisions in the report), cross-kind tag collisions, explicit/photo hook-finale-outro, A5 leak cap, map labels, face-aware scrapbook cover at exact size; photo-story loud serif fallback, POST tag options, per-shot minimum units in music mode, rotation-aware clip size. Synthetic tests: tests/test_vlog_photostory_fixes.py. Not re-run on the real demo footage.
  • 2026-10-05 demo round: longform-to-short sample-exact render (PCM joins, one AAC encode; |A−V| < 1 frame over 30 items), midpoint caption words at splits, per-episode hooks, text-density zoom centres, merged vertical manifests, pad_in word guard, split layout filling the 3:4 / 9:16 frame with a readable-text zoom (synthetic share frames); explainer compositions/ mkdir, display_en comma / “a million” numbers, per-sentence TTS check with re-takes, suffixed per-platform cover names, vertical fill rule + layout_check.py. Synthetic tests: tests/test_lfs_explainer_fixes.py. Not re-run on the real demo footage / a full HyperFrames render.
  • 2026-10-05 shared speech cleanup wiring: vlog fun speech shots (gentle by default, auto edits only, captions via cleanup.remap_words), vlog calm cleanup segments (analyze –ranges → apply before assembly), photo-story opt-in cleanup= / CLEANUP on audio="keep" clips, polish --cleanup pauses|gentle|standard|tight (off by default, never re-cuts a cleaned file). Synthetic tone-burst speech: tests/test_speech_cleanup_wiring.py. Not run on real footage / real whisper.