polish: make any exported edit publish-ready
Esta página aún no está traducida; aquí tienes la versión en inglés.
Use when: an edit has been exported from any NLE (Descript, CapCut/剪映, Premiere, Resolve, Final Cut, a
HyperFrames render…) and needs the last-mile treatment before upload: the first frame is dead (black title card,
fade-in, idle face) and should become the cover, the audio is too quiet (exports often land around -20 LUFS), the
pacing wants a small pitch-preserved speed bump, and the file should carry correct colour tags and faststart.
Inputs: export.mp4 (+ optional cover.png from workflows/cover). Outputs: final.mp4 at the source
resolution and frame rate, source-matched bitrate, persona.audio.loudness_lufs integrated loudness, bt709 tags.
With --platform <name>: the same, tuned to that platform’s profile + final.cover.jpg at the platform cover size;
with --platform a,b,...: the polished master plus one file + cover per platform and exports/manifest.json
(see Platforms).
Run from the project (video) folder. $VSTUDIO = repo root. Don’t polish a file you haven’t probed, and don’t
polish twice (a second loudness pass only costs headroom). Long-form uploads take the loudness and tag steps but
usually not the speed step. End-to-end context: references/SOP_SHORT_VIDEO.md.
Pipeline
Section titled “Pipeline”- Probe first, never blind-apply.
Terminal window ffprobe -v error -show_entries stream=codec_name,width,height,bit_rate,r_frame_rate -show_entries format=duration -of default=nw=1 export.mp4ffmpeg -y -ss 1 -i export.mp4 -frames:v 1 /tmp/check.png # ground-truth frame size (ffprobe metadata has lied before) - Decide the speed WITH the user (short-form parity). Default speed is 1.0, but for shorts (vertical
9:16 / 3:4 and under ~3 min, or any
--platformof douyin / tiktok / youtube-shorts / xiaohongshu) propose 1.2× - or personaspeed.bodyif the creator has set one - and ask before applying, e.g. “This is a 2:10 vertical talking-head; speed it up to 1.2× (pitch kept, ~1:48)? 1.0 / 1.1 / 1.2?”. The original short-form recipe ran everything at 1.2×; we keep that as the suggestion, not a silent default. Long horizontal videos stay at 1.0 unless asked. Anything above personaspeed.cjk_max_intelligible(1.4) prints a warning: dense Chinese speech gets hard to follow - say so if the user asks for more. - One command does the rest:
Terminal window python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 \--cover cover.png --speed 1.2 --keep work/polish --check # add --platform douyin etc.--coverreplaces the first--cover-sec(1.0) seconds of picture; audio runs uninterrupted, total length unchanged. Cover is fitted to the source frame (--cover-fit fill|fit).--speedtakes a number (1.2) or a persona key (hook,body,fast_body…). Omit for 1.0 (see step 2).- Loudness target defaults to
persona.audio.loudness_lufs(-14), or the--platformprofile’s; override with--lufs/--tp. TP -1.5 dBTP, LRA 11. --keep DIRkeeps01_cover.mp4,02_speed.mov,03_loud.mp4for stage-by-stage debugging.--checkwritesfinal.first.pngso you can eyeball that frame 0 is the cover.
- Verify (the script prints these; re-run any time):
Terminal window ffmpeg -i final.mp4 -af loudnorm=I=-14:print_format=summary -f null - 2>&1 | grep -E "Input Integrated|Input True Peak"
Optional: studio sound (--studio-sound [light|standard|strong], default off)
Section titled “Optional: studio sound (--studio-sound [light|standard|strong], default off)”For phone / laptop / room audio (Descript Studio Sound, 剪映 人声增强 / 降噪): denoise, dereverb, voice EQ, de-ess and a
gentle compressor before the cover / speed / loudness steps (vstudio.studiosound, local, picture stream-copied).
standard for most rooms, light for an already clean phone take, strong for fans / AC / echo. Listen first with
python -m vstudio.studiosound export.mp4 --ab work/ab --at 30 (level-matched before / after clips).
Optional: background music (--music calm|warm|bright|tech|story|<file>, default off)
Section titled “Optional: background music (--music calm|warm|bright|tech|story|<file>, default off)”A bed ducked under the voice (audio.mix_bed, --music-db -30 LUFS before ducking). Moods are built-in beds generated by
vstudio.music (free for any use); a file path uses her own track.
Optional: 气口 / filler / repeat cleanup of the export (--cleanup, default off)
Section titled “Optional: 气口 / filler / repeat cleanup of the export (--cleanup, default off)”An NLE export that still has long pauses, 嗯/呃 or repeats can take the shared cleanup tool (vstudio.cleanup,
references/CLEANUP.md) as step 0, before the cover:
python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 --cleanup gentle # 1st pass: auto edits# read final_review.md (待确认 items) to the creator, then:python3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 --cleanup gentle \ --cleanup-reply "确认 3,5 / 保留 7" [--cleanup-verify]- Profiles:
pauses(气口 only, no word edits),gentle,standard,tight. Onlyautoedits are cut until the creator replies;--cleanup-verifyre-transcribes the cut and stops if a content word was lost. - The EDL (
<out stem>.cleanup.json) and<out stem>_review.mdsit next to the output and are reused on re-runs (ids stay valid;--cleanup-freshre-analyses). Transcript:--cleanup-transcript(any shape) else ASR (--cleanup-lang). - Never cut twice: a file that is itself a cleanup output (has its
.cleanup.jsonsidecar) is not cut again, andcleanup.apply’s duration guard refuses an EDL made from a different-length file. Exports from Descript / CapCut / 剪映 where the creator already removed pauses usually wantpausesor nothing. - Duration afterwards = cleaned length ÷ speed (the self-check’s expectation follows it).
Platforms
Section titled “Platforms”Profiles live in lib/vstudio/platform.py (numbers and sources: references/PLATFORMS.md; persona overrides under
platforms.<name>). --platform takes name[:orientation] - xiaohongshu:vertical (1080x1440), xiaohongshu:full,
douyin, tiktok, youtube, youtube-shorts, bilibili[:horizontal|vertical] - or a comma list.
No --platform = the platform-neutral polish above, byte-for-byte the same commands as before.
no --platform |
one target (--platform youtube) |
several (--platform douyin,xiaohongshu:vertical) |
|
|---|---|---|---|
| loudness | persona audio.loudness_lufs, -1.5 dBTP |
profile loudness (LUFS/TP); --lufs/--tp still win |
master at the default; each export re-normalised to its profile |
| video encode | match source bitrate (or --crf) |
same, capped at profile encode.maxrate (bitrate mode: -b:v min(src, cap), maxrate <= cap; CRF mode: --crf else profile crf, + profile maxrate/bufsize) |
master as before; exports use profile CRF + maxrate (vstudio.export) |
| canvas | source | source - not reframed; a warning if the aspect differs | reframed per profile (--reframe-mode face, pad-blur fallback) |
| length | - | platform.check_length warnings (max / min / sweet spot) |
per export, in manifest.json |
| cover | first-second replacement only | + <out stem>.cover.jpg at platform.cover_size (from --cover, else a frame), .feed.jpg preview where the feed crops |
+ <platform>-<orientation>.cover.jpg per target |
| fps | source | source; warning above the profile max | converted to the profile default above max |
# one platform: polish + platform-sized cover + length checkpython3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o final.mp4 --cover cover.png --platform youtube-shorts
# several platforms: polish the master once, then one file + cover per platform + exports/manifest.jsonpython3 $VSTUDIO/workflows/polish/scripts/polish.py export.mp4 -o master.mp4 --cover cover.png \ --platform douyin,xiaohongshu:vertical,youtube --out-dir exports [--cues cues.json]
# same export step on an already-polished masterPYTHONPATH=$VSTUDIO/lib python3 -m vstudio.export master.mp4 --platforms douyin,xiaohongshu:vertical \ --out exports --cover cover.png [--cover cover-16x9.png] [--cues cues.json] [--title "..."]- Captions: an export from an NLE usually has captions burned in; those get cropped/covered on other canvases. For
several platforms, export a caption-free master from the editor and pass
--cues(cues.json / SRT) so each platform gets captions sized for its own caption box. - Single target with a different aspect (e.g. a vertical export with
--platform youtube): polish does NOT reframe silently. Use the multi-target form orpython -m vstudio.exportto get a reframed file. - Pass several
--coverimages tovstudio.export(e.g. 3:4 and 16:9); each platform takes the closest aspect. - Check
manifest.jsonwarnings(loudness, length, reframe fallback, captions that do not fit) before upload.
Rules & gotchas (from real runs)
Section titled “Rules & gotchas (from real runs)”- One stage per ffmpeg call. Cover + speed + loudnorm in one filter graph is miserable to debug; intermediates are cheap.
- Keep source resolution and frame rate. Never downscale “to save bandwidth”; platforms re-encode anyway.
- Match source bitrate when re-encoding (
--bitrate match, the default:-b:v src -maxrate 1.15x -bufsize 2x). Plain-crf 18on a ~28 Mbps phone/Descript export lands at 4–6 Mbps: fine for upload, not for an archive master.--crf Nis available when you do want size over fidelity. - Two-pass loudnorm (measure →
linear=trueapply with the measured values). Single-pass loudnorm runs in dynamic mode and audibly pumps speech. The loudness step (vstudio.audio.loudnorm_2pass) outputs 48 kHz stereo AAC (mono exports are up-mixed before measuring) and raises the LRA target to the measured LRA so loudnorm stays linear. Normalize once, at the end of the chain; don’t normalize again afterwards (platforms do their own pass and a second one only costs headroom).--skip-if-closeleaves gain alone within 1 LU. - Speed uses
atempo, neverasetrate(pitch stays put). Factors outside 0.5–2.0 are chained automatically. Above ~1.3× speech starts to sound artificial; personaspeed.cjk_max_intelligible(1.4) triggers a warning. - Speed before loudnorm. The speed step keeps audio as 24-bit PCM so the loudness pass is the only lossy audio encode.
- bt709 tags: encodes are tagged at encode time; stream-copied H.264/HEVC are re-tagged with the
h264_metadata/hevc_metadatabitstream filter (no re-encode). Untagged SDR files can shift colour on some players. HDR sources (iPhone HLG/Dolby Vision) need tone-mapping before this step, not just tags. - faststart moves the moov atom to the front so the first frame shows before the full download.
- The cover-replacement assumes the first second of the export carries no important picture (a typical dead title card). If the first second has a hook shot, prepend a cover instead (concat a 0.5–1 s still + delay audio) or pick a frame from the hook as the cover.
Self-check
Section titled “Self-check”- Output resolution and fps equal the source
- Video bitrate close to the source (or CRF chosen on purpose)
- Integrated loudness within ±1 LU of target, true peak ≤ -1.0 dBTP
- Frame 0 is the cover, not black
- Audio at t=0 is the source’s audio at t=0 (the cover replaces picture only; nothing shifted)
- Duration ≈ source (cleaned length with
--cleanup) ÷ speed (±0.5 s) - Shorts: 1.2× (or persona
speed.body) was proposed and the user’s answer applied - With
--platform: no length warnings you didn’t mention; cover jpg at the platform size; manifest warnings read