Aller au contenu

promo-recut

Cette page n’est pas encore traduite : voici la version anglaise.

Use when someone has a talking-head recording about something they made or found and wants a premium 16:9 short (and optionally 9:16 / 3:4 versions laid out per platform) for 小红书 / YouTube / B站. The talk is tight-cut, and the talking head slides into a split-screen next to 3D screenshot cards that scroll to whatever is being discussed. The edit adds highlighter sweeps, chips, keyword subtitles, punch-ins, a freeze-frame with a prompt card zooming out of a screenshot, and a zoom-through into a framed highlights montage. It finishes with an outro stamp, an end card, a chapter progress bar, a cover and the post copy.

Inputs → Outputs. input/talk.mp4 (talking head) + screenshots (PNG/JPG) + optional input/highlights.mp4

  • links → work/ (graded raw, body.mp4, outro.mp4, montage.mp4, layout.json), promo/ (HyperFrames project, index.html, subset fonts), promo-vertical/ (optional), my-promo.mp4 (-14 LUFS, BT.709, faststart), cover-4x3.jpg / cover-16x9.jpg / cover-3x4.jpg, post.md.

Everything content-specific lives in ONE project config: promo.config.yaml (or .json). It holds the KEEP spans, the creator’s cleanup reply (cut.reply), subtitles, cards and their highlight rows, chips, hold point, montage clips and labels, chapters, stamp and end card, the cover and the post. Start from $VSTUDIO/workflows/promo-recut/examples/promo.config.example.yaml, which documents every key. Taste (rates, loudness, brand colours, title length, tags) comes from the persona (persona.local.yaml).

Pipeline (run from the project dir; $VSTUDIO = repo root)

Section titled “Pipeline (run from the project dir; $VSTUDIO = repo root)”
my-promo/
promo.config.yaml input/talk.mp4 input/highlights.mp4 input/*.png
  1. Transcribe + find what to cut (去气口 / filler / 重复 / 口误: the shared vstudio.cleanup)
    Terminal window
    cp $VSTUDIO/workflows/promo-recut/examples/promo.config.example.yaml promo.config.yaml # then edit
    python3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yaml --suggest
    Set cut.body / cut.outro KEEP spans by sentence first. The first run writes work/audio.json, using vstudio.asr.transcribe (word timestamps; mlx_whisper on Apple Silicon, else faster_whisper, else OpenAI whisper-1 if OPENAI_API_KEY is set). --suggest snaps the KEEP spans word-safe and runs the same cleanup every speech workflow uses (python -m vstudio.cleanup analyze, references/CLEANUP.md) → work/cleanup.json + the review sheet work/cleanup_review.md: 待确认 CONFIRM (semantic fillers 然后/就是/那个 the audio isolates, interjections, restarts, re-takes, fillers whisper glued onto the next word — the old hidden-onset PATCH), 自动删 AUTO (嗯/呃/um/uh, stutters, clear repeats), 气口 (pauses squeezed per cut.profile, never deleted), 保留 KEEP (looks like a real word). Each row shows …before【removed】after… and why. The creator listens to the CONFIRM rows (ffplay -ss <t-0.5> -t 2 work/audio.wav) and answers, e.g. 确认 3,5,9 / 保留 7; put it in the config verbatim as cut.reply. Without a reply only AUTO rows are cut (never confirm every row unheard: on real footage that deletes real words). Old configs with cut.drop / cut.patch still work: they are translated to approvals of the matching cleanup edits (else word-safe editor cuts) and printed as config: lines.
  2. Tight cut + montage + layout
    Terminal window
    python3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yaml
    python3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yaml --verify
    This grades the raw once (grade, optional hdr_tonemap) and runs cleanup.apply per part on it (AUTO rows + cut.reply): word edits cut from the silence after the previous kept word to the silence before the next, video frame-exact, audio sample-exact with a 20 ms equal-power crossfade at every join (A/V cannot drift), then a two-pass loudnorm to persona audio.voice_lufs. Each part gets a sidecar work/<part>.cleanup.json. The step builds the highlights montage with baked 0.3 s internal crossfades (cut.xfade_assemble, plain acrossfade) and writes work/layout.json: durations, raw→cut maps as vstudio.cut.TimeMap items, and word times. --verify = cleanup.verify per part: re-ASR, content words that went missing are printed with their raw / cut time (exit 1: add 保留 N for the edit covering it to cut.reply and re-cut); leftover fillers / repeats are listed. Changing the KEEP spans renumbers the edits, so with a reply in the config the cut refuses until --suggest is re-run and the reply re-confirmed. Also listen to every join before going on.
  3. Subtitles + cards
    Terminal window
    python3 $VSTUDIO/workflows/promo-recut/scripts/tight_cut.py promo.config.yaml --draft-subs # paste, then edit
    python3 $VSTUDIO/workflows/promo-recut/scripts/find_rows.py input/shot-post.png --preview work/rows.png
    Write subtitles in raw seconds and wrap the key term in 【】. find_rows.py measures text-row bands with numpy. Use those y0/y1 rows for cards[].highlights, box and scroll. Its width_frac column is a good starting highlight width.
  4. Build the HyperFrames project
    Terminal window
    python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml
    python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --orientation vertical # optional
    cd promo && npx hyperframes lint && npx hyperframes snapshot --at <split-in>,<hold>,<zoom-through>,<outro> --no-end
    This writes index.html and timeline.json, copies media into assets/, and extracts the freeze frame. It also subsets Noto Sans SC + STIX Two Text from vstudio.config.FONT_DIR (vstudio.render.subset_project_fonts; --no-fonts reuses assets/fonts/) to just the characters in the config (≈60 KB per CJK weight). Look at snapshots taken mid-transition, not only mid-scene. Expect 0 lint errors. The nested_structure_needs_subcomposition warnings are advisory.
  5. Render + deliver
    Terminal window
    bash $VSTUDIO/workflows/promo-recut/scripts/export.sh promo my-promo.mp4 delivery # = python3 .../export.py
    This renders with HyperFrames, runs a two-pass loudnorm to persona audio.loudness_lufs (-14; audio.loudnorm_2pass, --lufs overrides), writes BT.709 colour tags into the H.264/HEVC stream without re-encoding and adds +faststart (media.retag_bt709). It then prints the streams and the measured loudness. Use --skip-render <raw.mp4> <out.mp4> to redo only the delivery step.
  6. Cover + post
    Terminal window
    python3 $VSTUDIO/workflows/promo-recut/scripts/make_cover.py promo.config.yaml
    python3 $VSTUDIO/workflows/promo-recut/scripts/post_copy.py promo.config.yaml
    The cover is vstudio.cover.split_cover fed from the config. It takes a frame from the talk (or a given image) and retouches it through vstudio.retouch (slim, eye, de-shine, skin, light makeup, optional body slim). It crops around the detected face. Landscape sizes get a split cover: photo on one side, and on the other a dark panel with quote, title + highlighted term, a framed highlights thumbnail, chips, a red tag and a 记笔记 tag. Portrait sizes stack the photo on top. The post (vstudio.publish.post_body) gets the title (length checked per platform), body, links, a chapter timeline from promo/timeline.json (MM:SS, floored) and tags.

Timeline model (what build_promo computes)

Section titled “Timeline model (what build_promo computes)”

body plays at rates.body and is split at hold.at into body + freeze image + body2 (data-media-start). These share track 2 back to back. The montage starts zoom_through s (0.5) before the body ends and runs on its own track 3 at rates.montage. The outro starts where the montage’s nominal length ends, while the montage clip keeps running another zoom_through s underneath. The end card follows. Raw second → final second is BT(raw) = TimeMap.to_final(raw, snap="fwd")/rate (+ hold if raw ≥ hold.at), so you never type a final-timeline time. Chapter anchors are raw body seconds or start / montage / outro.

  • Overlapping clips need separate tracks and a known audio owner. Body and montage overlap during the zoom-through, so they sit on different data-track-index values. Only overlap where the outgoing clip is silent (the cut adds about 0.1 s of tail after the last word). If an overlap would play two voices, set data-volume="0" on one, or use a muted copy, rather than relying on a fade.
  • The outgoing scene must stay alive during a transition. body2 runs until M + TZ, and the montage runs TZ past the outro start. If a clip ends exactly when its exit animation starts, the transition shows the plate (a black flash).
  • Overlay start states must be opacity: 0 in CSS (gsap_fullscreen_overlay_starts_visible). This covers the dim layer, prompt box, screen, outro wrap, badges, stamp and end-card lines. Otherwise seek-based rendering shows them on frame 0 or before their fromTo runs.
  • Chapter labels need a scrim. Small labels over bright footage are unreadable. The bar sits on a bottom gradient (#bar-scrim), and labels get a text shadow.
  • Whisper merges fillers into neighbouring words. A “word” that is too long for its characters usually starts with a hidden 然后/嗯. The cleanup finds the energy dip and lists it as a 粘连口头禅 (filler-merged) CONFIRM row that cuts only up to the rise, so the word itself stays; confirm it by ear.
  • ASR the cut to verify it (--verify). Joins that look right on the word list can still swallow a syllable or keep half a filler.
  • Captions must clear before a zoom-through. build_promo clips any cue that crosses M + TZ so it ends at M - 0.08. A caption flying into the montage frame looks broken.
  • Measure screenshot rows, don’t guess. Highlight bands and red boxes use image-pixel rows from find_rows.py (row-band detection on the grey-level difference from the background). Card scale is card_width / image_width, and build_promo applies it.
  • Adjacent card windows closer than 0.8 s are merged into one split, so the face doesn’t bounce back to full frame between cards.
  • Montage clips with label null are transitional fragments. The previous step label continues over them.
  • Fonts: only Noto Sans SC / STIX Two Text (OFL), subset per video. A system CJK font that exists only on your machine silently falls back in the headless renderer.

build_promo.py has a GEO table for horizontal (card box, split inset, subtitle line, bar, screen frame, label positions); override any value with layout.horizontal.*. The vertical geometry is not a table: it is computed from a platform profile (vertical_geo, next section), so every element lands inside that platform’s safe box and clear of its button column. Override any computed key with layout.vertical.*.

Terminal window
python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml # horizontal (unchanged legacy layout)
python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --orientation vertical # persona platforms.default at 9:16 -> promo-vertical/
python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --platform xiaohongshu:full # -> promo-vertical-xiaohongshu-full/
python3 $VSTUDIO/workflows/promo-recut/scripts/build_promo.py promo.config.yaml --platform douyin # -> promo-vertical-douyin-vertical/

--platform (or config platform:) takes any vstudio.platform profile: xiaohongshu:full (9:16), xiaohongshu:vertical (3:4, 1080x1440), douyin, tiktok, youtube-shorts, bilibili:vertical; a horizontal profile keeps the 1920x1080 layout. --out DIR names the project dir. Per vertical profile the layout is:

element where (from platform.safe_box / caption_box / keepouts)
chapter bar + labels top of the safe box, on a top scrim (labels 22 px)
talking head (split) a band under the bar, ~42 % of the free height (~2.4 face heights when a face is detected), full safe width; the clip window is centred on the speaker’s face (vstudio.face on 6 frames, else centred; layout.vertical.face: [fx, fy]), and object-position keeps the face centred when full-frame
chips + screenshot card under the band down to 24 px above the caption band; side margin widened to clear the lower-right button column
captions the profile’s caption box, bottom-anchored, text-wrap: balance, size = box width / max_chars_zh clamped to the profile’s caption size range; cues that can’t fit 2 lines are warned about
prompt hold card the card column, mid-height
montage screen + step label / badge / tag 16:9 screen at the safe width, centred in the free area but kept above the button column
stamp / end card stamp top-left of the free area; end card centred between the safe top and the caption band

timeline.json records platform, canvas and the boxes used (safe, caption, keep-outs, face band, card, screen) so snapshots can be checked against them. The build also warns when the length is outside the profile’s sweet spot / max and when chapter labels collide on the narrower bar (merge or shorten chapters). Check every vertical build: npx hyperframes lint (0 errors) and snapshots at a card, the hold, mid zoom-through, the montage and the outro.

Multi-platform delivery. Build one project per canvas, render each, then hand the renders to vstudio.export, which (for the same aspect) only scales, re-loudnorms to the profile’s LUFS / true peak, encodes with the profile’s settings, crops covers and writes post stubs + manifest.json:

Terminal window
python3 $VSTUDIO/workflows/promo-recut/scripts/export.py promo my-promo.mp4 --platform youtube # or plain persona LUFS
python3 -m vstudio.export my-promo.mp4 --platforms youtube,bilibili,xiaohongshu:horizontal --out exports/ \
--cover cover-16x9.jpg --cover cover-4x3.jpg --post post.json
python3 $VSTUDIO/workflows/promo-recut/scripts/export.py promo-vertical-douyin-vertical my-promo-douyin.mp4 --platform douyin
python3 -m vstudio.export my-promo-douyin.mp4 --platforms douyin,tiktok,youtube-shorts --out exports/ --cover cover-3x4.jpg

Captions here are part of the design (keyword highlight, timed with the cards), so they are burned in by HyperFrames per canvas. If you want vstudio.export to place them instead, build with --clean-master (no burned subtitles) and pass --cues <promo_dir>/cues.json (final-timeline cues, written by every build). Do not reframe the 16:9 promo to 9:16 with vstudio.export: the split screen and cards would be cropped. Build the vertical layout instead.

Every packaging effect here (split screen, screenshot cards with scroll / highlighter / red box, chips, keyword subtitles, punch-ins, freeze hold, zoom-through into a framed screen, step labels, badge, tag, title, stamp, end card) is a generator in lib/vstudio/hf.py returning {css, html, js}. build_promo only lays out times and geometry and passes hf.JS("D.X") references into its const D data object. To reuse one effect in another HyperFrames project, call the generator with plain numbers. Catalogue: references/EFFECTS.md.