longform-to-short: long recording → cut course video and/or N short episodes
Esta página aún no está traducida; aquí tienes la versión en inglés.
Use when: someone hands over a long recording — a lecture, webinar, livestream replay, podcast with screen-share, a Zoom / Meet / Teams teaching or coaching session (typically 30–120 min) — and wants it “剪成课程 / 上课实录 / 教学长视频”, “切片发小红书 / 分几集”, wants dead air and off-topic chatter removed, browser chrome / bookmark bar cropped off, chapters + hook + zoom + 记笔记 panels + subtitles added, a participant’s voice anonymised, covers and 发布文案. Also for tweaking any one of those on an already-cut video (every decision is a config value, so it is a cheap re-run).
Inputs: one recording (raw/my-talk.mp4), optionally its embedded caption track (speaker
labels), optionally a live web demo to re-record. Outputs (in out/): final_subbed.mp4
(1080p, hook cold-open, chapter cards, body sped up, code zoom cut-ins, note panels, burned subs
with 勘误 notes, ~−14 LUFS), subs.srt (soft CC), cover_16x9.png, cover_3x4.png,
episodes/ep{N}.mp4 + ep{N}_cover.png, 发布包.md, platform_checks.json; with vertical targets also
vertical/ep{N}/<platform>-<orientation>.mp4 (3:4 / 9:16 slices: speaker or title band on top, the
shared screen following the content below, captions in the platform’s caption box) + covers, post stubs and
vertical/manifest.json.
Not in scope: writing a script (the recording exists), music beds (tested and rejected for lectures — only a card stinger).
export VSTUDIO=/path/to/video-studio # repo root; ./install.sh fetched fonts + whispercd my-project # your video foldermkdir -p raw work out && mv ~/somewhere/my-talk.mp4 raw/cp $VSTUDIO/workflows/longform-to-short/examples/config.example.py work/config.pyEdit work/config.py as you go (set src first). Every script is
python3 $VSTUDIO/workflows/longform-to-short/scripts/<step>.py work/config.py; below $S means
$VSTUDIO/workflows/longform-to-short/scripts. Scripts run inside work/ (all JSON / PNG / segment
scratch lives there). Taste (speeds, loudness, accent colour, tags, term fixes, title limits) comes
from persona.local.yaml; per-video decisions from the config. Needs ffmpeg (with libass for
the burn step — else pip install static-ffmpeg and it is used automatically), Pillow, numpy,
scipy, and mlx-whisper (Apple Silicon) or faster-whisper.
Pipeline
Section titled “Pipeline”1. Analyse the source
python3 $S/analyze.py work/config.py # audio16k.wav, rec_subs.srt, sounded.json, geo/*.pngpython3 $S/transcribe.py work/config.py # -> audio16k.json (minutes; run in background; cached)python3 $S/geometry.py work/config.py # -> geometry.json + crop_spans.json (chrome removed)python3 $S/speaker_timeline.py work/config.py # -> speakers.json (labels + times only)Set speakers.aliases so real display names become roles (“Host”, “Guest”) — unknown labels are
anonymised as “Speaker A/B…”. Look at one geo/ frame and set share (x0,y0,x1,y1 of the main
shared page) and source_size.
2. Review blocks → keep / drop (human or LLM reads the transcript)
python3 $S/make_blocks.py work/config.py # -> blocks.txt (numbered, speaker-tagged)Read blocks.txt; fill keep.ranges + keep.chapters. For a public video: drop chatter, logistics,
participant-specific advice, anything that names or identifies a participant; keep the teaching.
python3 $S/build_keep_list.py work/config.py # -> keep_list.json + cleanup_review.ep<N>.md, prints chapter tableSpeech cleanup (气口 / filler / 重复 / 口误) is the shared tool vstudio.cleanup (repo
references/CLEANUP.md), not workflow code. Per kept segment: its edges go through cleanup.snap_range
(word-safe, keep.pad_in/pad_out never reach a dropped neighbour word), its end through
cleanup.extend_end (the 40 ms cut fade lands after the last word’s real tail), cuts split it into
sub-ranges snapped the same way, and cleanup.clean runs per sub-range (profile cleanup.profile, else
persona cleanup.profile, else standard; cleanup.overrides, cleanup.enabled: false to skip). The kept
pieces go into keep_list.json (keep); build_timeline makes one clip per piece.
Only AUTO edits are applied at first. One sheet per episode (episodes.items chapter ranges; one sheet
without items): work/cleanup_review.ep<N>.md lists 待确认 / 自动删 / 气口 / 保留 with context. The creator
replies per sheet → cleanup.reply: {"2": "确认 3,5 / 保留 7"} in the config (or
build_keep_list.py --reply 2:"确认 3,5 / 保留 7"), re-run from build_keep_list. Hook clips (hook.src,
episodes.items[].hook.src) are snapped + end-extended at the hook speed → work/hook_edges.json.
3. Polish targets
python3 $S/transient_scan.py work/config.py # accidental tab / desktop flashes -> transients.jsonpython3 $S/zoom_targets.py work/config.py # focal centre for each zoom.windows entryEyeball each transient; real ones → freezes (the scan is visual - tab / desktop flashes - which the speech
cleanup cannot see). Participant questions to keep → pitches.windows
(decide by content). Fillers / pauses / repeats are already cleaned (step 2); off-topic asides or anything
the sheet missed → cuts (word-safe; re-run build_keep_list). Pick the hook clip → hook.src + hook.lines.
For slice jobs give each episode its own cold open instead: episodes.items[i].hook: {src: [t0, t1], lines: [..]}
(placed right before that episode’s first chapter card; the episode starts at it; item 0’s replaces hook.src).
Zoom windows are hand-written: put zoom.windows [[t0, t1], ...] (source s) where the speaker walks
through code / a doc passage (from the transcript + geo/ frames). zoom_targets.py only finds each
window’s centre: a grey code block (zoom.code_rgb) if there is one, else the densest text in a
zoom.box-sized window (GitHub / Notion / plain white pages, dark editors); zoom.centers overrides by hand.
4. (optional) Re-record a stale live demo
python3 $S/probe_app.py work/config.py # screenshot + selectors of demo.urlpython3 $S/record_demo.py work/config.py # Playwright video -> set demo.rec, demo.enabled: TrueNarration audio is kept; only the video of the demo.keep_idx segments is swapped.
5. Master edit decision
python3 $S/build_timeline.py work/config.py # -> timeline.json (cards, clips, freezes, pitch, zoom)6. Overlays + render
python3 $S/make_assets.py work/config.py # cards/card_NN.png + hook_overlay.pngpython3 $S/make_panels.py work/config.py # panels/*.png + panels.jsonpython3 $S/make_audio_assets.py work/config.py # card_sting.wavpython3 $S/render.py work/config.py # -> out/final.mp4 (loudnorm to persona LUFS)7. Subtitles + final burn
python3 $S/build_subs.py work/config.py # subs.srt + subs.ass (term fixes, 勘误 notes)python3 $S/burn_final.py work/config.py # panels + subs in one encode -> out/final_subbed.mp48. Covers, episodes, publish package
python3 $S/make_cover.py work/config.py # out/cover_16x9.png + out/cover_3x4.pngpython3 $S/make_episodes.py work/config.py # out/episodes/* + out/发布包.mdepisodes.count: N auto-splits at chapter-card boundaries; episodes.items gives explicit chapter
ranges + per-episode cover copy. With neither, only the full-video 发布包 is written. For a
“short episodes only” job, still build the full cut and just publish the episodes.
Every episode is checked against the episode targets (length sweet spot / max, title) and the full cut
against the long-form targets → out/platform_checks.json + WARN lines (see Platforms).
8c. (optional) Vertical slices — see “Vertical slices” below
python3 $S/make_vertical.py work/config.py # out/vertical/ep{N}/<platform>-<orientation>.mp4 + manifestpython3 $S/scan_popups.py out/vertical/ep1/xiaohongshu-vertical.mp4 --plan work/vertical/1080x1440/plan.json # editor popups still visible (2 fps; exit 1 + times when found) # + menus / toolbars left open (static overlays; --no-static)make_vertical runs the same 2 fps scans on every master (plan.json visible_popups, static_overlays: a
floating card - faint border, drop shadow - that stays open over whole items, which the popup scan misses).
9. Verify before handing off
python3 $S/qa.py work/config.py # decode, loudness, mosaics, pitch check- decode is clean; integrated loudness ≈ persona
audio.loudness_lufs(−14). - Open every
work/qa/mosaic_NN.jpg: zero participant avatars, name tags, bookmark bars, emails. - Pitch windows dropped ~3 semitones; the host’s voice unchanged.
- Spot-check each cut join (re-transcribe ~6 s of final audio around it or just listen).
Platforms
Section titled “Platforms”targets (list) or platform (one value / comma list) in the config, or --targets / --platform on
make_cover.py, make_episodes.py, make_vertical.py: values like youtube, bilibili,
xiaohongshu:horizontal, xiaohongshu:vertical (3:4), xiaohongshu:full (9:16), douyin, tiktok,
youtube-shorts (python -m vstudio.platform lists them all with their boxes). Profiles come from
lib/vstudio/platform.py + persona platforms.<name> overrides (references/PLATFORMS.md in the repo root).
| What | Unset (default) | With explicit targets |
|---|---|---|
| target list | youtube + <persona platforms.default>:horizontal |
as given |
16:9 canvas (render.py) |
render.size or 1920x1080 |
first horizontal target’s canvas unless render.size is set |
loudness (render.py) |
persona audio.loudness_lufs, −1.5 dBTP |
that profile’s loudness (lufs / tp) |
burned captions (build_subs.py) |
22 CJK chars/line, ASS default margin | max_chars_zh per line, bottom margin from caption_box (subtitles.max_line still wins) |
记笔记 panel (burn_final.py) |
panel_pos |
panel_pos clamped into safe_box |
covers (make_cover.py) |
16:9 + 3:4 | + every other aspect a target needs (cover_9x16.png for 抖音 / TikTok / Shorts) |
| episode covers | ep{N}_cover.png (3:4) |
+ ep{N}_cover_<aspect>.png per episode-target aspect |
| length / title checks | full cut vs long-form targets (sweet spot > 5 min: YouTube, B站), episodes vs short-form ones (小红书, 抖音, TikTok, Shorts); episodes.targets overrides → out/platform_checks.json |
same |
| vertical slices | make_vertical.py uses <persona platforms.default>:vertical |
every vertical target |
The 16:9 outputs are byte-for-byte the pre-platform ones when targets / platform is unset (checked on the
synthetic run). episodes.max_minutes (15) still warns too.
Clean master + multi-platform export. build_subs.py writes work/cues.json (final-time cues) and,
per horizontal target, work/cues.<platform>-<orientation>.json re-laid to fit that profile’s caption box;
burn_final.py --clean-master writes out/master_clean.mp4 (cards, hook, panels, NO captions). Then one
command makes a per-platform file with captions placed and sized for each UI, loudness to each target,
length / title warnings, covers re-fitted to each platform’s size and a post stub:
python3 $S/burn_final.py work/config.py --clean-masterpython3 -m vstudio.export out/master_clean.mp4 --platforms youtube --cues work/cues.youtube-horizontal.json \ --cover out/cover_16x9.png --post work/post.json --out out/exports # one call per caption trackFor 3:4 / 9:16 don’t export the 16:9 master (a face reframe or pad-blur of a screen share is unreadable and
can reveal tiles): use make_vertical.py, which renders a vertical master from the source and then calls the
same vstudio.export per episode and target. python -m vstudio.export on
work/vertical/<W>x<H>/master.mp4 with --cues work/vertical/cues.<platform>-<orientation>.json --start/--dur re-exports
one slice by hand.
Vertical slices
Section titled “Vertical slices”make_vertical.py turns the cut into 3:4 / 9:16 slices for every vertical target (one master per canvas,
then per episode and target vstudio.export). The source is the original recording, so the layout can use
regions the 16:9 render cropped away. Layout per item: --mode > vertical.segments [{src: [t0, t1], mode}]
(source s) > episodes.items[i].vertical > vertical.mode (default split):
- split (default; the classic lecture slice): speaker band on top (40 %), screen below, running to the
frame bottom (
split.screen_to: frame; the part under the caption box is dimmed bysplit.scrimso captions stay readable, reading start / activity are kept in the part above it). The speaker band is avstudio.reframeface-mode crop ofvertical.speaker.region(the host’s cam tile, small tiles upscaled for detection). If there is no region, or faces are found in <speaker.min_hit(0.3) of the frames, the band becomes a title band (16 %,split.band_frac): series + current chapter title, the hook lines during the hook. A camera-less screen share therefore never shows participant tiles or avatar name tags. - screen: the whole content area (to the frame bottom, like split) is the screen. The crop is the main
text block (column of ink; panel borders ignored), zoomed until a text line is ≥
screen.min_text_px(28) tall on the canvas (line height measured on a full-res frame), withinscreen.min_scale..max_scale(1.6–3.0x of the source); a share region too short to fill the box atmax_scaleis drawn at the top of the box and the leftover sits under the captions / platform UI, centred when the block fits, else line starts kept visible. y follows on-screen activity inside the block (frame differences at 4 Hz: typing, highlights, new blocks; cursor-sized changes ignored), code-zoom windows use theirzoom_windows.jsoncentre, a page switch cuts to the first text block. The path goes throughvstudio.reframe.follow(vstudio.filters.OneEuro + dead zone + eased, speed-limited pan). - speaker: speaker crop only. pad-blur: the whole screen region fitted to the width over a blurred fill.
Layout boxes come from the profiles on that canvas (most conservative of them): content = safe-box top to
caption-box top, captions are re-laid per profile (work/vertical/cues.<platform>-<orientation>.json: long cues split so each fits
max_lines at a size inside the profile’s range, ≤ max_chars_zh per line; size capped when two lines would
not fit the caption box height, e.g. 60 px on 小红书 3:4). Chapter cards are re-drawn at the vertical canvas;
记笔记 panels are re-drawn ≤ 62 % of the safe width, right-aligned at the top of the screen area.
Privacy. The screen crop never leaves the item’s crop (geometry span = shared page minus browser chrome /
bookmark bar, or the zoom box); the speaker crop never leaves speaker.region; vertical.exclude rects
(participant tiles, name tags) are painted out of every source frame first; work/vertical/<W>x<H>/plan.json
reports privacy_overlap_frames (must be 0). Only set speaker.region to the HOST’s own camera.
Outputs: out/vertical/ep{N}/<platform>-<orientation>.mp4 (+ .cover.jpg from cover_3x4.png /
cover_9x16.png, .post.md, .crop.json, manifest.json), out/vertical/manifest.json (canvas, measured
loudness, length / title / label warnings; merged across runs, so --targets youtube-shorts:vertical after a
小红书 run adds entries instead of replacing them). qa.py summarises them and writes contact sheets.
Speed: ~2x real time per canvas on a laptop (decode, compose in numpy, x264), plus one export per target.
Hard-won defaults
Section titled “Hard-won defaults”- Platform auto-captions are unusable for non-English speech (Meet/Zoom map Chinese onto random English words). Keep only speaker labels + timestamps; re-transcribe with whisper.
- Whisper flags:
condition_on_previous_text=False,hallucination_silence_threshold=2, else it loops on 嗯嗯嗯 over silence (transcribe.py→vstudio.asr.transcribesets both). Put domain terms inasr_prompt(or--prompt);--backend mlx|faster|openaiforces an engine. Results are cached inwork/audio16k.wav.asr.json, so re-runs are instant. - Speaker labels are often misattributed. Decide who said what by content before cutting or pitch-shifting anyone.
- Cut from the transcript, not from silence. Meeting audio is too noisy for silencedetect-driven
cutting;
sounded.jsononly estimates dead air. - Sparse keyframes: fast input-seek +
-frames:v 1can land seconds off in meeting recordings. Stills seek ~25 s early and decode forward (vstudio.media.grab_frame); the transient scan decodes the whole file at 1 fps. - Overlaying a PNG needs
-loop 1+overlay=…:shortest=1, or the 1-frame stream ends before itsenablewindow and silently never renders. - libass: some ffmpeg builds (e.g. recent Homebrew) ship without the
assfilter;burn_final.pyasksvstudio.media.ffmpeg_bin(need=["ass"]), which falls back tostatic_ffmpeg. - VideoToolbox: use
-q:vquality mode, never a fixed-b:v(balloons the file). For code / slide text libx264 crf ~20 is crisper (burn.encoder: x264). - Never stream-copy-concat AAC segments. Each AAC segment carries encoder priming / padding, so a
-c copyconcat drifts ~25 ms per join (0.7 s over a 4-min cut: later episodes opened with the previous one’s last words).render.pyrenders each item to its exact frame / sample count on one grid (_lfc.segment_grid), joins PCM, and encodes AAC once;make_vertical.pyuses the same grid. - Captions at splits: a word straddling a zoom / pitch split goes to the piece holding its midpoint
(
build_subs.py), so “面试” never comes out as “试”. - Seamless zoom cut-ins: zoom/pitch/freeze splits are marked audio-continuous, so no afade at
those joins; fades only at real cuts (30 ms in / 40 ms out) - every cleanup /
cutsjoin is one. - Many pieces: cleanup turns a segment into several clips (one per kept piece), and render.py encodes each
item separately: a long lecture renders slower than before;
--from Nstill reuses earlier segments. - No BGM. A synthesized hook bed was tried and rejected as odd under a lecture; only a 1.6 s card stinger. Don’t propose music beds for this genre.
- Geometry thresholds assume a light page on a dark meeting canvas; dark-mode shares need
geometry.bright/transients.lumaretuned.
Publishing notes
Section titled “Publishing notes”- YouTube: full upload; timestamps in the description auto-create chapters (first must be 00:00);
upload
subs.srtas CC; 16:9 cover. - 小红书: regular-video cap is ~15 min (verify in the live uploader) → split into episodes at card
boundaries; post them as 3:4 vertical slices (
targets: [..., "xiaohongshu:vertical"]) or 16:9. Chapter labels ≤ 14 chars (publish.short_labels); titles checked withxhs_lenagainst personaplatforms.xiaohongshu.title_max. Portrait 3:4 cover (feeds show a 4:3 crop). - B站: optional archive.
- If the lecture contains a now-outdated claim, keep the burned 勘误 note (
subtitles.errata) and pinpublish.errata_lineas a comment.
Iterating
Section titled “Iterating”Change one knob, re-run from that step down:
- cleanup reply / profile,
cuts, keep ranges, hook src →build_keep_list→build_timeline→ … as below. - freezes / pitches / speeds →
build_timeline→render(--from Nreuses earlier segments) →build_subs→burn_final. - panel text only →
make_panels→burn_final(no body re-render). - term fixes / errata →
build_subs→burn_final. - cover copy →
make_cover; episode copy →make_episodes --no-video. - vertical layout / targets →
make_vertical(--master-onlyto look at the masters first,--episodes 2to re-export one); anything upstream (cuts, subs, panels) → re-runmake_verticalafter it.
Coaching-call variant (diarization, visual privacy leaks, word-level filler tightening):
references/coaching-variant.md.