Skip to content

longform-to-short: long recording → cut course video and/or N short episodes

Use when: someone hands over a long recording — a lecture, webinar, livestream replay, podcast with screen-share, a Zoom / Meet / Teams teaching or coaching session (typically 30–120 min) — and wants it “剪成课程 / 上课实录 / 教学长视频”, “切片发小红书 / 分几集”, wants dead air and off-topic chatter removed, browser chrome / bookmark bar cropped off, chapters + hook + zoom + 记笔记 panels + subtitles added, a participant’s voice anonymised, covers and 发布文案. Also for tweaking any one of those on an already-cut video (every decision is a config value, so it is a cheap re-run).

Inputs: one recording (raw/my-talk.mp4), optionally its embedded caption track (speaker labels), optionally a live web demo to re-record. Outputs (in out/): final_subbed.mp4 (1080p, hook cold-open, chapter cards, body sped up, code zoom cut-ins, note panels, burned subs with 勘误 notes, ~−14 LUFS), subs.srt (soft CC), cover_16x9.png, cover_3x4.png, episodes/ep{N}.mp4 + ep{N}_cover.png, 发布包.md, platform_checks.json; with vertical targets also vertical/ep{N}/<platform>-<orientation>.mp4 (3:4 / 9:16 slices: speaker or title band on top, the shared screen following the content below, captions in the platform’s caption box) + covers, post stubs and vertical/manifest.json.

Not in scope: writing a script (the recording exists), music beds (tested and rejected for lectures — only a card stinger).

Terminal window
export VSTUDIO=/path/to/video-studio # repo root; ./install.sh fetched fonts + whisper
cd my-project # your video folder
mkdir -p raw work out && mv ~/somewhere/my-talk.mp4 raw/
cp $VSTUDIO/workflows/longform-to-short/examples/config.example.py work/config.py

Edit work/config.py as you go (set src first). Every script is python3 $VSTUDIO/workflows/longform-to-short/scripts/<step>.py work/config.py; below $S means $VSTUDIO/workflows/longform-to-short/scripts. Scripts run inside work/ (all JSON / PNG / segment scratch lives there). Taste (speeds, loudness, accent colour, tags, term fixes, title limits) comes from persona.local.yaml; per-video decisions from the config. Needs ffmpeg (with libass for the burn step — else pip install static-ffmpeg and it is used automatically), Pillow, numpy, scipy, and mlx-whisper (Apple Silicon) or faster-whisper.

1. Analyse the source

Terminal window
python3 $S/analyze.py work/config.py # audio16k.wav, rec_subs.srt, sounded.json, geo/*.png
python3 $S/transcribe.py work/config.py # -> audio16k.json (minutes; run in background; cached)
python3 $S/geometry.py work/config.py # -> geometry.json + crop_spans.json (chrome removed)
python3 $S/speaker_timeline.py work/config.py # -> speakers.json (labels + times only)

Set speakers.aliases so real display names become roles (“Host”, “Guest”) — unknown labels are anonymised as “Speaker A/B…”. Look at one geo/ frame and set share (x0,y0,x1,y1 of the main shared page) and source_size.

2. Review blocks → keep / drop (human or LLM reads the transcript)

Terminal window
python3 $S/make_blocks.py work/config.py # -> blocks.txt (numbered, speaker-tagged)

Read blocks.txt; fill keep.ranges + keep.chapters. For a public video: drop chatter, logistics, participant-specific advice, anything that names or identifies a participant; keep the teaching.

Terminal window
python3 $S/build_keep_list.py work/config.py # -> keep_list.json + cleanup_review.ep<N>.md, prints chapter table

Speech cleanup (气口 / filler / 重复 / 口误) is the shared tool vstudio.cleanup (repo references/CLEANUP.md), not workflow code. Per kept segment: its edges go through cleanup.snap_range (word-safe, keep.pad_in/pad_out never reach a dropped neighbour word), its end through cleanup.extend_end (the 40 ms cut fade lands after the last word’s real tail), cuts split it into sub-ranges snapped the same way, and cleanup.clean runs per sub-range (profile cleanup.profile, else persona cleanup.profile, else standard; cleanup.overrides, cleanup.enabled: false to skip). The kept pieces go into keep_list.json (keep); build_timeline makes one clip per piece. Only AUTO edits are applied at first. One sheet per episode (episodes.items chapter ranges; one sheet without items): work/cleanup_review.ep<N>.md lists 待确认 / 自动删 / 气口 / 保留 with context. The creator replies per sheet → cleanup.reply: {"2": "确认 3,5 / 保留 7"} in the config (or build_keep_list.py --reply 2:"确认 3,5 / 保留 7"), re-run from build_keep_list. Hook clips (hook.src, episodes.items[].hook.src) are snapped + end-extended at the hook speed → work/hook_edges.json.

3. Polish targets

Terminal window
python3 $S/transient_scan.py work/config.py # accidental tab / desktop flashes -> transients.json
python3 $S/zoom_targets.py work/config.py # focal centre for each zoom.windows entry

Eyeball each transient; real ones → freezes (the scan is visual - tab / desktop flashes - which the speech cleanup cannot see). Participant questions to keep → pitches.windows (decide by content). Fillers / pauses / repeats are already cleaned (step 2); off-topic asides or anything the sheet missed → cuts (word-safe; re-run build_keep_list). Pick the hook clip → hook.src + hook.lines. For slice jobs give each episode its own cold open instead: episodes.items[i].hook: {src: [t0, t1], lines: [..]} (placed right before that episode’s first chapter card; the episode starts at it; item 0’s replaces hook.src). Zoom windows are hand-written: put zoom.windows [[t0, t1], ...] (source s) where the speaker walks through code / a doc passage (from the transcript + geo/ frames). zoom_targets.py only finds each window’s centre: a grey code block (zoom.code_rgb) if there is one, else the densest text in a zoom.box-sized window (GitHub / Notion / plain white pages, dark editors); zoom.centers overrides by hand.

4. (optional) Re-record a stale live demo

Terminal window
python3 $S/probe_app.py work/config.py # screenshot + selectors of demo.url
python3 $S/record_demo.py work/config.py # Playwright video -> set demo.rec, demo.enabled: True

Narration audio is kept; only the video of the demo.keep_idx segments is swapped.

5. Master edit decision

Terminal window
python3 $S/build_timeline.py work/config.py # -> timeline.json (cards, clips, freezes, pitch, zoom)

6. Overlays + render

Terminal window
python3 $S/make_assets.py work/config.py # cards/card_NN.png + hook_overlay.png
python3 $S/make_panels.py work/config.py # panels/*.png + panels.json
python3 $S/make_audio_assets.py work/config.py # card_sting.wav
python3 $S/render.py work/config.py # -> out/final.mp4 (loudnorm to persona LUFS)

7. Subtitles + final burn

Terminal window
python3 $S/build_subs.py work/config.py # subs.srt + subs.ass (term fixes, 勘误 notes)
python3 $S/burn_final.py work/config.py # panels + subs in one encode -> out/final_subbed.mp4

8. Covers, episodes, publish package

Terminal window
python3 $S/make_cover.py work/config.py # out/cover_16x9.png + out/cover_3x4.png
python3 $S/make_episodes.py work/config.py # out/episodes/* + out/发布包.md

episodes.count: N auto-splits at chapter-card boundaries; episodes.items gives explicit chapter ranges + per-episode cover copy. With neither, only the full-video 发布包 is written. For a “short episodes only” job, still build the full cut and just publish the episodes. Every episode is checked against the episode targets (length sweet spot / max, title) and the full cut against the long-form targets → out/platform_checks.json + WARN lines (see Platforms).

8c. (optional) Vertical slices — see “Vertical slices” below

Terminal window
python3 $S/make_vertical.py work/config.py # out/vertical/ep{N}/<platform>-<orientation>.mp4 + manifest
python3 $S/scan_popups.py out/vertical/ep1/xiaohongshu-vertical.mp4 --plan work/vertical/1080x1440/plan.json
# editor popups still visible (2 fps; exit 1 + times when found)
# + menus / toolbars left open (static overlays; --no-static)

make_vertical runs the same 2 fps scans on every master (plan.json visible_popups, static_overlays: a floating card - faint border, drop shadow - that stays open over whole items, which the popup scan misses).

9. Verify before handing off

Terminal window
python3 $S/qa.py work/config.py # decode, loudness, mosaics, pitch check
  • decode is clean; integrated loudness ≈ persona audio.loudness_lufs (−14).
  • Open every work/qa/mosaic_NN.jpg: zero participant avatars, name tags, bookmark bars, emails.
  • Pitch windows dropped ~3 semitones; the host’s voice unchanged.
  • Spot-check each cut join (re-transcribe ~6 s of final audio around it or just listen).

targets (list) or platform (one value / comma list) in the config, or --targets / --platform on make_cover.py, make_episodes.py, make_vertical.py: values like youtube, bilibili, xiaohongshu:horizontal, xiaohongshu:vertical (3:4), xiaohongshu:full (9:16), douyin, tiktok, youtube-shorts (python -m vstudio.platform lists them all with their boxes). Profiles come from lib/vstudio/platform.py + persona platforms.<name> overrides (references/PLATFORMS.md in the repo root).

What Unset (default) With explicit targets
target list youtube + <persona platforms.default>:horizontal as given
16:9 canvas (render.py) render.size or 1920x1080 first horizontal target’s canvas unless render.size is set
loudness (render.py) persona audio.loudness_lufs, −1.5 dBTP that profile’s loudness (lufs / tp)
burned captions (build_subs.py) 22 CJK chars/line, ASS default margin max_chars_zh per line, bottom margin from caption_box (subtitles.max_line still wins)
记笔记 panel (burn_final.py) panel_pos panel_pos clamped into safe_box
covers (make_cover.py) 16:9 + 3:4 + every other aspect a target needs (cover_9x16.png for 抖音 / TikTok / Shorts)
episode covers ep{N}_cover.png (3:4) + ep{N}_cover_<aspect>.png per episode-target aspect
length / title checks full cut vs long-form targets (sweet spot > 5 min: YouTube, B站), episodes vs short-form ones (小红书, 抖音, TikTok, Shorts); episodes.targets overrides → out/platform_checks.json same
vertical slices make_vertical.py uses <persona platforms.default>:vertical every vertical target

The 16:9 outputs are byte-for-byte the pre-platform ones when targets / platform is unset (checked on the synthetic run). episodes.max_minutes (15) still warns too.

Clean master + multi-platform export. build_subs.py writes work/cues.json (final-time cues) and, per horizontal target, work/cues.<platform>-<orientation>.json re-laid to fit that profile’s caption box; burn_final.py --clean-master writes out/master_clean.mp4 (cards, hook, panels, NO captions). Then one command makes a per-platform file with captions placed and sized for each UI, loudness to each target, length / title warnings, covers re-fitted to each platform’s size and a post stub:

Terminal window
python3 $S/burn_final.py work/config.py --clean-master
python3 -m vstudio.export out/master_clean.mp4 --platforms youtube --cues work/cues.youtube-horizontal.json \
--cover out/cover_16x9.png --post work/post.json --out out/exports # one call per caption track

For 3:4 / 9:16 don’t export the 16:9 master (a face reframe or pad-blur of a screen share is unreadable and can reveal tiles): use make_vertical.py, which renders a vertical master from the source and then calls the same vstudio.export per episode and target. python -m vstudio.export on work/vertical/<W>x<H>/master.mp4 with --cues work/vertical/cues.<platform>-<orientation>.json --start/--dur re-exports one slice by hand.

make_vertical.py turns the cut into 3:4 / 9:16 slices for every vertical target (one master per canvas, then per episode and target vstudio.export). The source is the original recording, so the layout can use regions the 16:9 render cropped away. Layout per item: --mode > vertical.segments [{src: [t0, t1], mode}] (source s) > episodes.items[i].vertical > vertical.mode (default split):

  • split (default; the classic lecture slice): speaker band on top (40 %), screen below, running to the frame bottom (split.screen_to: frame; the part under the caption box is dimmed by split.scrim so captions stay readable, reading start / activity are kept in the part above it). The speaker band is a vstudio.reframe face-mode crop of vertical.speaker.region (the host’s cam tile, small tiles upscaled for detection). If there is no region, or faces are found in < speaker.min_hit (0.3) of the frames, the band becomes a title band (16 %, split.band_frac): series + current chapter title, the hook lines during the hook. A camera-less screen share therefore never shows participant tiles or avatar name tags.
  • screen: the whole content area (to the frame bottom, like split) is the screen. The crop is the main text block (column of ink; panel borders ignored), zoomed until a text line is ≥ screen.min_text_px (28) tall on the canvas (line height measured on a full-res frame), within screen.min_scale..max_scale (1.6–3.0x of the source); a share region too short to fill the box at max_scale is drawn at the top of the box and the leftover sits under the captions / platform UI, centred when the block fits, else line starts kept visible. y follows on-screen activity inside the block (frame differences at 4 Hz: typing, highlights, new blocks; cursor-sized changes ignored), code-zoom windows use their zoom_windows.json centre, a page switch cuts to the first text block. The path goes through vstudio.reframe.follow (vstudio.filters.OneEuro + dead zone + eased, speed-limited pan).
  • speaker: speaker crop only. pad-blur: the whole screen region fitted to the width over a blurred fill.

Layout boxes come from the profiles on that canvas (most conservative of them): content = safe-box top to caption-box top, captions are re-laid per profile (work/vertical/cues.<platform>-<orientation>.json: long cues split so each fits max_lines at a size inside the profile’s range, ≤ max_chars_zh per line; size capped when two lines would not fit the caption box height, e.g. 60 px on 小红书 3:4). Chapter cards are re-drawn at the vertical canvas; 记笔记 panels are re-drawn ≤ 62 % of the safe width, right-aligned at the top of the screen area.

Privacy. The screen crop never leaves the item’s crop (geometry span = shared page minus browser chrome / bookmark bar, or the zoom box); the speaker crop never leaves speaker.region; vertical.exclude rects (participant tiles, name tags) are painted out of every source frame first; work/vertical/<W>x<H>/plan.json reports privacy_overlap_frames (must be 0). Only set speaker.region to the HOST’s own camera.

Outputs: out/vertical/ep{N}/<platform>-<orientation>.mp4 (+ .cover.jpg from cover_3x4.png / cover_9x16.png, .post.md, .crop.json, manifest.json), out/vertical/manifest.json (canvas, measured loudness, length / title / label warnings; merged across runs, so --targets youtube-shorts:vertical after a 小红书 run adds entries instead of replacing them). qa.py summarises them and writes contact sheets. Speed: ~2x real time per canvas on a laptop (decode, compose in numpy, x264), plus one export per target.

  • Platform auto-captions are unusable for non-English speech (Meet/Zoom map Chinese onto random English words). Keep only speaker labels + timestamps; re-transcribe with whisper.
  • Whisper flags: condition_on_previous_text=False, hallucination_silence_threshold=2, else it loops on 嗯嗯嗯 over silence (transcribe.py → vstudio.asr.transcribe sets both). Put domain terms in asr_prompt (or --prompt); --backend mlx|faster|openai forces an engine. Results are cached in work/audio16k.wav.asr.json, so re-runs are instant.
  • Speaker labels are often misattributed. Decide who said what by content before cutting or pitch-shifting anyone.
  • Cut from the transcript, not from silence. Meeting audio is too noisy for silencedetect-driven cutting; sounded.json only estimates dead air.
  • Sparse keyframes: fast input-seek + -frames:v 1 can land seconds off in meeting recordings. Stills seek ~25 s early and decode forward (vstudio.media.grab_frame); the transient scan decodes the whole file at 1 fps.
  • Overlaying a PNG needs -loop 1 + overlay=…:shortest=1, or the 1-frame stream ends before its enable window and silently never renders.
  • libass: some ffmpeg builds (e.g. recent Homebrew) ship without the ass filter; burn_final.py asks vstudio.media.ffmpeg_bin(need=["ass"]), which falls back to static_ffmpeg.
  • VideoToolbox: use -q:v quality mode, never a fixed -b:v (balloons the file). For code / slide text libx264 crf ~20 is crisper (burn.encoder: x264).
  • Never stream-copy-concat AAC segments. Each AAC segment carries encoder priming / padding, so a -c copy concat drifts ~25 ms per join (0.7 s over a 4-min cut: later episodes opened with the previous one’s last words). render.py renders each item to its exact frame / sample count on one grid (_lfc.segment_grid), joins PCM, and encodes AAC once; make_vertical.py uses the same grid.
  • Captions at splits: a word straddling a zoom / pitch split goes to the piece holding its midpoint (build_subs.py), so “面试” never comes out as “试”.
  • Seamless zoom cut-ins: zoom/pitch/freeze splits are marked audio-continuous, so no afade at those joins; fades only at real cuts (30 ms in / 40 ms out) - every cleanup / cuts join is one.
  • Many pieces: cleanup turns a segment into several clips (one per kept piece), and render.py encodes each item separately: a long lecture renders slower than before; --from N still reuses earlier segments.
  • No BGM. A synthesized hook bed was tried and rejected as odd under a lecture; only a 1.6 s card stinger. Don’t propose music beds for this genre.
  • Geometry thresholds assume a light page on a dark meeting canvas; dark-mode shares need geometry.bright / transients.luma retuned.
  • YouTube: full upload; timestamps in the description auto-create chapters (first must be 00:00); upload subs.srt as CC; 16:9 cover.
  • 小红书: regular-video cap is ~15 min (verify in the live uploader) → split into episodes at card boundaries; post them as 3:4 vertical slices (targets: [..., "xiaohongshu:vertical"]) or 16:9. Chapter labels ≤ 14 chars (publish.short_labels); titles checked with xhs_len against persona platforms.xiaohongshu.title_max. Portrait 3:4 cover (feeds show a 4:3 crop).
  • B站: optional archive.
  • If the lecture contains a now-outdated claim, keep the burned 勘误 note (subtitles.errata) and pin publish.errata_line as a comment.

Change one knob, re-run from that step down:

  • cleanup reply / profile, cuts, keep ranges, hook src → build_keep_list → build_timeline → … as below.
  • freezes / pitches / speeds → build_timeline → render (--from N reuses earlier segments) → build_subs → burn_final.
  • panel text only → make_panels → burn_final (no body re-render).
  • term fixes / errata → build_subs → burn_final.
  • cover copy → make_cover; episode copy → make_episodes --no-video.
  • vertical layout / targets → make_vertical (--master-only to look at the masters first, --episodes 2 to re-export one); anything upstream (cuts, subs, panels) → re-run make_vertical after it.

Coaching-call variant (diarization, visual privacy leaks, word-level filler tightening): references/coaching-variant.md.