Output edit: second-pass editing of every finished clip (`python -m vstudio.project output`)
Every finished video the desk app shows can be edited again (成片二次编辑): a recipe project’s export
(<item>/<platform>-<orientation>) or a video of an adopted work folder (final/A.mp4; a plain skill folder is
adopted on first use). Code: lib/vstudio/project/outputs.py (outputs, edit document, ops, AI),
outrender.py (cached render), outfx.py (effect catalogue + frame layers). Tests: tests/test_output_edit.py, tests/test_output_cut.py (transcript cuts).
1. Modes and capability flags
Section titled “1. Modes and capability flags”| mode | when | what a render does |
|---|---|---|
pipeline |
our pipeline kept the clip’s clean (caption-free) master + cues: talkinghead compose, or a work folder’s .vstudio/masters.json (outputs.register_master(dir, output, master, cues)) / <stem>.master.mp4 / <stem>_clean.mp4 next to it or in work/; the master must match the output’s duration (±0.35 s) |
re-composes from the master: captions are ours (text, style, position, keyword colour, on / off), other aspects are a fresh face-tracked reframe of the master |
flattened |
only the finished file | edits go on top of it: trims / cuts / speed re-encode, effects and a title band overlay it; captions can only be ADDED, over a mask (blur / solid box hiding the burned ones, detected band by default) or in a band layout (whole frame scaled into the middle, title band top, caption band bottom); burned text is never restyled; other aspects reframe the file (layout band keeps the whole frame) |
show returns caps (mode, trim, cut, cut_snap, speed, loudness, effects, title_band, cover, export, undo, ai, captions_ours, caption_text, caption_style, caption_restyle_burned, caption_toggle, caption_add, caption_placements, relayout, relayout_layouts, burned_captions, audio, cut_words, cut_strategy) and caps_notes
(why a flag is off). cut_words: true = cuts by transcript word index (cut {words} / cut {gap}, section 3a);
cut_strategy says what a transcript cut does to the picture: remaster (pipeline: captions are re-composed from
the master and re-timed in the same step), snap_captions (flattened with burned captions detected: cut edges snap to
the burned caption line boundaries), hard (flattened, no burned captions detected or unknown: a plain cut with a
2-frame audio fade at every join; no video fade - a video dip on a hard cut reads as a flash).
2. Files (nothing else in the project / work folder is written; the original output is never overwritten)
Section titled “2. Files (nothing else in the project / work folder is written; the original output is never overwritten)”<project>/state/outputs/<slug>/ or <work>/.vstudio/outputs/<slug>/ edit.json {version, output, mode, source_sig, steps [{id, at, by user|ai, note, ops, describe, revert_of?}], redo, burned} chat.json {version, output, turns [{id, at, role user|ai, text, context, proposed, dropped, summary, provider, model, cost_usd, seconds, status draft|applied|discarded|reverted|note, applied_step, reverted_by}]} transcript.json word timings of the output (cut snapping, ai context), keyed by the file signature cache/<stage>-<key>.mp4|wav|jpg stage results (4 kept per stage) renders/<target>.<preview|final>.mp4, <target>.cover.jpg, manifest.json {renders {target.quality: {file, key, at}}} .vstudio/status.json heartbeats of this output's rendersAudio-only edits are fast. When the edit leaves the picture of the output’s own canvas as it came (no trim /
cut / speed / grade / end fade / transition / frame effect, captions / title / theme unchanged) and only adds audio
work (studio-sound 人声增强, music-bed 背景音乐, sfx-placement, loudness), the final stage is remux: the
original file’s video stream is copied and only the new audio is encoded (outrender.picture_untouched). A 2:28
1080x1920 clip: 473 s (frame pass) -> 11 s.
3. Ops (edit --ops '[{...}]'; one call = one undo step; all or nothing)
Section titled “3. Ops (edit --ops '[{...}]'; one call = one undo step; all or nothing)”Times are seconds on the ORIGINAL output timeline ("time_base": "edited" converts from the current edited
timeline); positions are fractions of the canvas (x, y = centre).
| op | fields | notes |
|---|---|---|
trim |
start?, end?, snap=true |
word-safe edges (cleanup.snap_range) when a transcript is cached |
cut / cut_remove |
start, end, why?, no_snap? or words: [i0, i1], sig?, why? or gap: i, keep?, sig? / index |
snapped to whole words with cleanup.snap_cut on the output’s transcript (transcribed once, cached), then each edge moves to the nearest audio zero crossing (±5 ms); transcript cuts: section 3a |
speed |
value 0.5-2.5 |
setpts + atempo |
loudness |
lufs, tp? |
default: the target platform’s profile |
captions |
enabled |
pipeline only |
caption_text |
cue, text, force? |
pipeline cues must stay faithful (proofread.faithful; not-faithful unless force); added cues (a<n>) any text |
caption_style |
`style {size 0.6-1.8, color, highlight, keywords [..], stroke, stroke_color, position bottom | middle |
caption_add / caption_remove |
start, end, text or from_transcript: true, start?, end? / cue |
flattened: placement auto = mask over the detected caption band |
caption_placement |
`mode band | mask |
title |
text, sub?, color?, band_color?, y?, height?, size? |
text: "" removes the band |
theme |
`theme editorial | mono |
effect_add |
effect, start, end? / duration?, params {}, id? |
effect id / alias / zh label from the catalogue; params validated (unknown dropped, clamped, required checked) |
effect_update |
id, start?, end?, shift?, duration?, params? |
move / retime / change params |
effect_remove |
id |
|
cover |
`t, text, style card | plain |
export_add / export_remove |
target (platform[:orientation], 3:4, 9:16, 16:9), `layout auto |
band` |
reset |
everything back to the original (undoable) |
3a. Transcript cuts (Descript-style editing)
Section titled “3a. Transcript cuts (Descript-style editing)”Word indices refer to the cached transcript (transcript.json words [{w, t, te, p?}], paths.transcript in
show); show.words_sig is its signature (sha1 of [w, t, te] per word, 10 hex; null while nothing is
transcribed - show never transcribes).
| op form | meaning |
|---|---|
{op: cut, words: [i0, i1], sig?, why: transcript} |
cut words i0..i1 (inclusive): start = W[i0].t, end = W[i1].te, then snap_cut (word-safe edges; if the audio refinement would drop a selected word the plain word times are used). sig != the current words_sig -> stale-words (nothing applied); indices not ints / out of range / reversed -> bad-param |
{op: cut, gap: i, keep: 0.25, sig?, why: pause} |
shorten the pause between word i and i+1, keeping keep s (0.05-2, default 0.25) of it: cut [W[i].te + keep/2, W[i+1].t - keep/2], no word snapping; a pause shorter than keep + 0.05 -> bad-param |
{op: cut, start, end, why: filler} |
a plain range (marks, AI) - snapped as before |
Every cut edge (unless no_snap, or no audio) then moves to the nearest zero crossing of the audio the renderer cuts
(the master in pipeline mode) within ±5 ms (outputs.zero_cross(samples, sr, t, t0) on ~20 ms decoded with one
tiny ffmpeg call per edge; any failure leaves the edge): values[i].zero_cross: true. Cut times keep 5 decimals.
Flattened + snap_captions: each edge of a transcript cut (words or why: transcript) moves to the nearest
burned caption line boundary within 0.3 s (lines approximated by subs.cues_from_words(W, max_chars=14));
an edge that cannot snap keeps its place and adds warning cut-splits-burned-caption {at}
(values[i].caption_snap true / false).
why is free text; the desk sends transcript, filler, pause. values[i] of a cut: {cut [a, b], words (what was said in it), snapped, kept_s, word_range?, gap?, gap_s?, keep_s?, zero_cross?, caption_snap?}; its history line
is op-cut {start, end, why, said}.
One Apply = one step, captions and effects re-timed in it. When a step holds a cut with why transcript /
filler / pause (or a words / gap cut), derived ops are computed against the final state of the step and
appended to the SAME step (a single undo restores everything):
- captions (pipeline cues): a cue with a word cut by this step loses the cut words (
caption_text {cue, text: join_words(remaining), force: true}); nothing left (or no edited span) ->caption_remove; under 2 characters left -> removed and its text appended to the previous kept cue. Added captions are only re-mapped by the render. - effects (visual, overlapping a new cut): edited length < 0.4 s or fully cut ->
effect_remove;pop-wordswhose word (itstextamong the words in its window, else the word at its start) was cut ->effect_remove; partly cut -> nothing (the render maps it), reported as trimmed.sfx-placementstarting inside a new cut ->effect_remove.music-bed, joins, grade, end fade and the progress bar are never touched. step.retimed(also onhistory.steps[].retimed):{captions {retimed, shortened, removed}, effects {trimmed [{id, effect, label {en, zh}, from_s, to_s}], removed [{id, effect, label}]}, sfx_removed, targets [primary, ...exports]}- every version updates: the cut happens before the canvas work of every target. Derived ops describe asop-caption-text {cue, text, force},op-caption-remove {cue},op-effect-remove {id}.
Marks (show.marks, from the cached transcript only, at most 500, never fails show):
[{kind filler|pause|lowconf, i0, i1, text, save_s, group}] - fillers = cleanup.detect without audio, rows of
kind filler not kept (group = the normalized filler, e.g. 那个; save_s = its length); pauses = gaps over 0.6 s
between words i0 and i1 = i0 + 1 (text “1.2s”, save_s = gap - 0.25, group pause; cut it with {gap: i0});
lowconf = words with an ASR probability p < 0.5 (when the transcript has it).
preview-edl (output preview-edl --project P --output O --ops JSON --json, outputs.preview_edl(d, output, ops)): what the renderer would keep if the draft ops were applied; nothing is written (the transcript is read
from the cache). Ops are normalized + folded in order exactly like edit (same snapping); a failing op is skipped
and reported. -> {ok, output, keep [[a, b]] (outputs.segments, the renderer's own), cuts [[a, b]], duration (edited s), source_duration, ops [{index, op (normalized), describe, value}], dropped [{index, op, error}], warnings}. No derived caption / effect ops (they never change the kept ranges).
ai: output ai --instruction "..." [--apply] (or edit --op ai): the routed model (task output_edit of
vstudio.llm; claude-code / codex / API / local routes all work) gets the output facts, caps, current state,
the op list, the effect catalogue and the captions / transcript, and answers {ops, summary}. Every proposed op
goes through the same validator as edit; invented effects / ops / params, out-of-range values and ops the caps
forbid are dropped with their error. Without --apply nothing changes (the desk shows proposed for confirmation,
then calls edit --ops with ops, or ai --apply). No model configured: a literal-phrase fallback (speed, head /
tail trim, platform exports, LUFS, progress bar, fade) with warning no-model.
Selective revert (output revert --step <id>): cancels ONE earlier step and keeps every step after it. It is
recorded as a new step {op: revert, step} (revert_of on the step; undo / redo it like any step; reverting a
revert brings its target back); the state is folded without the cancelled step (history.steps[].reverted).
Refused with revert-conflict {steps, n} when a later step builds on it (edits / removes an effect or added caption
it created, removes a cut by index after it changed the cut list, or a later reset): the honest alternative is
undo back to it. Also unknown-step, already-reverted.
Context (ai --context '{"range": [12, 18], "cues": ["c4"], "effect": "fx2"}'): what the creator points at
(timeline selection, caption cues, one effect instance). Validated (bad-context, unknown-effect-instance), passed
to the model as focus (cue text and the effect instance spelled out) with “this / here / 这段 mean the focus”;
the literal fallback understands “剪掉这段 / cut this” and “zoom” on the range. Echoed as context.
Chat transcript (chat.json, output chat): every ai call records a turn (--no-record to skip) with its
proposals, provider, model, cost and seconds, status draft; edit --turn ID marks it applied with the step
id; revert of that step marks it reverted. The desk adds its own turns (slash-command cards) with
chat --add JSON and patches status with chat --turn ID --set JSON. show returns chat so the conversation
is rebuilt after a reopen.
Project-level AI (python -m vstudio.project ai --project P --instruction T [--outputs all|a,b] [--context JSON] [--timeout 120] [--json | --json-events], projai.py): one request for every output (“remove the series
label from all clips”). (1) read every targeted output (cached transcripts only); (2) a cheap rule check BEFORE any
model call: removing / replacing text burned into a flattened output, or restyling burned captions, can never be
done by the editor, so those outputs get needs_rerender at once (no model, < 2 s) - evidence needed: a text noun
(标题 / 水印 / label …), the text in the folder’s own scripts, or a cached transcript that does not say it (what
is SAID goes to the model as a cut); text this editor added (title band, an effect’s text, an added caption) is
removed with plain ops; (3) the rest goes to ONE model call (task output_edit) with every output’s summary,
answered as groups per output, each op validated against that output’s caps. --timeout bounds each provider; a
timed-out / failed provider falls back along the route’s chain (fallback event + fallback in the result).
Nothing is applied: the desk applies a group with output edit (one undo step per output).
needs_rerender = {outputs, titles, code burned-text | burned-captions-restyle | model-needs-rerender, reason, targets, paths [{kind rerender-scripts {files [{file, line, text}], scripts, prompt} | regenerate {items, command} | re-export, message, actions [{kind copy-prompt {prompt} | open-file {file, line} | regenerate {items} | reveal {file}, label}]}]}. --json-events: {event: stage, stage: read|check|ask|plan, n, provider?, elapsed},
{event: partial, needs_rerender, groups} (the rule answer, sent BEFORE the model call so the desk shows it at
once), {event: fallback, from, to, code, error}, then {event: done, result} (or {event: failed, code, ...},
exit 5). A failed model call keeps the rule answer (warning llm-failed, failed {provider, code}).
4. Effects (output effects --json [--no-thumbs])
Section titled “4. Effects (output effects --json [--no-thumbs])”Each id is a row of the effects registry (vstudio.effects); {id, label {en, zh}, description {en, zh}, kind, stage, params (JSON-Schema subset with x-zh), required, default_dur, aliases, thumbnail}; thumbnails are sample
frames rendered once into $VSTUDIO_CACHE/output_fx/.
| id | zh | stage | params |
|---|---|---|---|
pop-words |
弹出大字 | frame | text, color, size, angle, anim, x, y |
stacking-stamps |
印章 | frame | text, angle, scale, anim, x, y |
punch-in |
推镜放大 | frame | scale, in_dur, out_dur, hold return |
quote-card |
金句卡 | frame | text, sub, speaker, width, theme, anim, x, y |
callout-bubble |
标注气泡 / 箭头 | frame | text, arrow_x, arrow_y, theme, anim, x, y |
chapter-card |
章节卡 | frame | title, index, total, accent |
notes-panel |
记笔记面板 | frame | title, bullets, theme, width, anim, x, y |
overlay-images |
贴纸 / 标签 | frame | image, text, style tag |
badge |
角标 | frame | text, color, scale, anim, x, y |
red-box |
框选高亮 | frame | x, y, w, h, color, width, anim |
progress-bar-pil |
进度条 | frame | style line |
marker-sweep |
马克笔划重点 | frame | text (【kw】), size, surface auto |
chapter-rule |
章节细线 | frame | label, title, index, width, anim, x, y |
number-counter |
数字滚动 | frame | value, label, prefix, suffix, count, size, anim, x, y |
lower-third |
人名条 | frame | name, role, scale, anim, x, y |
sfx-placement |
音效 | audio | name (the synthesized bank), gain |
music-bed |
背景音乐 | audio | file, duck_db, music_lufs (whole clip) |
xfade-joins |
转场 | timeline | transition (any vstudio.xfade name), duration: at a cut join = an xfade, elsewhere a flash / dip |
end-fade |
结尾淡出 | timeline | duration |
vlog-grade |
调色 | timeline | sat, contrast, warm |
5. Render (output render [--quality preview|final] [--targets primary,douyin:vertical|all] [--json-events])
Section titled “5. Render (output render [--quality preview|final] [--targets primary,douyin:vertical|all] [--json-events])”Stages per target: canvas (reframe the master / file to the target canvas) -> timeline (trim, cuts, joins,
speed, grade, end fade) -> audio (SFX, music bed, two-pass loudness) -> final (frame pass: mask / band layout,
punch-in, overlays, cards, title band, captions, progress bar; encode) -> cover. Each stage key hashes its input
key + only the ops it uses, so a caption / overlay change re-runs final only, an SFX change audio + final, a
cut everything from timeline; an unchanged render (or one undone back to) is all cache hits. preview = same
resolution, low bitrate (ultrafast, CRF 30, 2 Mb/s); final = the platform delivery encode. show.renders says
which renders are fresh for the current state. Every render writes .vstudio/status.json heartbeats
(output-edit:<stage>, progress during the frame pass) into the edit folder and the owning project / work
folder; the owner’s previous record is restored afterwards.
6. JSON contract (desk app)
Section titled “6. JSON contract (desk app)”All commands take --json (paths absolute). Every user-facing message (caps notes, warnings, refusals, op
descriptions) is {code, params, message (English), message_zh}: the desk localises by code + params; content
(captions, titles, cover text) stays in the project’s content language.
| command | stdout |
|---|---|
output list --project P |
`{ok, dir, kind project |
output show --project P --output O |
{ok, output {id, file, mode, canvas, fps, duration, platform, master}, caps, caps_notes, state, captions [{id, start, end, text, original, edited, removed, added}], effects [{id, effect, start, end, params, label, edited [a, b]}], timeline {segments, joins, speed, duration}, history {steps [{id, at, by, note, describe, retimed}], undo, redo}, renders [{target, quality, file, fresh}], warnings, paths {doc, dir, renders, transcript}, words_sig, marks} |
output edit ... --ops JSON |
the show document + {step {id, ops, describe, retimed?}, values [per op], warnings} |
output preview-edl ... --ops JSON |
{ok, output, keep, cuts, duration, source_duration, ops, dropped, warnings} (nothing written) |
output ai ... --instruction T [--context JSON] [--apply] |
{ok, context, proposed [{op, normalized, describe, why, warnings}], ops, dropped [{op, error}], summary, provider, model, cost_usd, seconds, warnings, turn, applied, step?} |
output render ... |
{ok, output, mode, quality, targets [{target, file, cover, canvas, layout, duration, key, cached, stages [{stage, key, cached, seconds}], warnings}], seconds}; --json-events: target-start, stage-done, target-done, render-done lines; --with-ops JSON: a before / after preview of ops that are not applied, into renders/<target>.compare.mp4 (compare: true; edit.json and the manifest untouched) |
output revert ... --step ID |
the show document + {step, reverted} |
output chat ... [--add JSON / --turn ID --set JSON] |
{ok, turns} / {ok, turn} |
| `output undo | redo …` |
output effects |
{ok, effects [...], n} |
ai --project P --instruction T [--outputs] |
`{ok, scope project, answer changes |
Errors: exit 5 + {ok: false, error, code, params, message, message_zh}. Codes: unknown-owner,
unknown-output, unknown-op, bad-op, no-ops, unknown-effect, unknown-effect-instance,
duplicate-effect, bad-param, bad-time, too-short, cut-no-word, unknown-cut, no-audio, no-words,
transcribe-failed, captions-not-ours, unknown-cue, empty-text, not-faithful, placement-pipeline,
unknown-target, nothing-to-undo, nothing-to-redo, unknown-step, already-reverted, revert-conflict,
bad-context, unknown-turn, llm-failed, llm-bad-json, render-failed, stale-words {expected, got}.
Warnings / notes: flattened, captions-add-only, relayout-crops-burned, no-master, master-mismatch,
source-changed, placement-auto, style-added-only, band-pipeline, effect-cut-away, param-adjusted,
unknown-param, no-model, length, cut-splits-burned-caption {at}. Op descriptions: op-<op> (e.g. op-effect-add {effect, start, end},
op-revert {step, what}).