Aller au contenu

Output edit: second-pass editing of every finished clip (`python -m vstudio.project output`)

Cette page n’est pas encore traduite : voici la version anglaise.

Every finished video the desk app shows can be edited again (成片二次编辑): a recipe project’s export (<item>/<platform>-<orientation>) or a video of an adopted work folder (final/A.mp4; a plain skill folder is adopted on first use). Code: lib/vstudio/project/outputs.py (outputs, edit document, ops, AI), outrender.py (cached render), outfx.py (effect catalogue + frame layers). Tests: tests/test_output_edit.py, tests/test_output_cut.py (transcript cuts).

mode when what a render does
pipeline our pipeline kept the clip’s clean (caption-free) master + cues: talkinghead compose, or a work folder’s .vstudio/masters.json (outputs.register_master(dir, output, master, cues)) / <stem>.master.mp4 / <stem>_clean.mp4 next to it or in work/; the master must match the output’s duration (±0.35 s) re-composes from the master: captions are ours (text, style, position, keyword colour, on / off), other aspects are a fresh face-tracked reframe of the master
flattened only the finished file edits go on top of it: trims / cuts / speed re-encode, effects and a title band overlay it; captions can only be ADDED, over a mask (blur / solid box hiding the burned ones, detected band by default) or in a band layout (whole frame scaled into the middle, title band top, caption band bottom); burned text is never restyled; other aspects reframe the file (layout band keeps the whole frame)

show returns caps (mode, trim, cut, cut_snap, speed, loudness, effects, title_band, cover, export, undo, ai, captions_ours, caption_text, caption_style, caption_restyle_burned, caption_toggle, caption_add, caption_placements, relayout, relayout_layouts, burned_captions, audio, cut_words, cut_strategy) and caps_notes (why a flag is off). cut_words: true = cuts by transcript word index (cut {words} / cut {gap}, section 3a); cut_strategy says what a transcript cut does to the picture: remaster (pipeline: captions are re-composed from the master and re-timed in the same step), snap_captions (flattened with burned captions detected: cut edges snap to the burned caption line boundaries), hard (flattened, no burned captions detected or unknown: a plain cut with a 2-frame audio fade at every join; no video fade - a video dip on a hard cut reads as a flash).

2. Files (nothing else in the project / work folder is written; the original output is never overwritten)

Section titled “2. Files (nothing else in the project / work folder is written; the original output is never overwritten)”
<project>/state/outputs/<slug>/ or <work>/.vstudio/outputs/<slug>/
edit.json {version, output, mode, source_sig, steps [{id, at, by user|ai, note, ops, describe, revert_of?}], redo, burned}
chat.json {version, output, turns [{id, at, role user|ai, text, context, proposed, dropped, summary, provider,
model, cost_usd, seconds, status draft|applied|discarded|reverted|note, applied_step, reverted_by}]}
transcript.json word timings of the output (cut snapping, ai context), keyed by the file signature
cache/<stage>-<key>.mp4|wav|jpg stage results (4 kept per stage)
renders/<target>.<preview|final>.mp4, <target>.cover.jpg, manifest.json {renders {target.quality: {file, key, at}}}
.vstudio/status.json heartbeats of this output's renders

Audio-only edits are fast. When the edit leaves the picture of the output’s own canvas as it came (no trim / cut / speed / grade / end fade / transition / frame effect, captions / title / theme unchanged) and only adds audio work (studio-sound 人声增强, music-bed 背景音乐, sfx-placement, loudness), the final stage is remux: the original file’s video stream is copied and only the new audio is encoded (outrender.picture_untouched). A 2:28 1080x1920 clip: 473 s (frame pass) -> 11 s.

3. Ops (edit --ops '[{...}]'; one call = one undo step; all or nothing)

Section titled “3. Ops (edit --ops '[{...}]'; one call = one undo step; all or nothing)”

Times are seconds on the ORIGINAL output timeline ("time_base": "edited" converts from the current edited timeline); positions are fractions of the canvas (x, y = centre).

op fields notes
trim start?, end?, snap=true word-safe edges (cleanup.snap_range) when a transcript is cached
cut / cut_remove start, end, why?, no_snap? or words: [i0, i1], sig?, why? or gap: i, keep?, sig? / index snapped to whole words with cleanup.snap_cut on the output’s transcript (transcribed once, cached), then each edge moves to the nearest audio zero crossing (±5 ms); transcript cuts: section 3a
speed value 0.5-2.5 setpts + atempo
loudness lufs, tp? default: the target platform’s profile
captions enabled pipeline only
caption_text cue, text, force? pipeline cues must stay faithful (proofread.faithful; not-faithful unless force); added cues (a<n>) any text
caption_style `style {size 0.6-1.8, color, highlight, keywords [..], stroke, stroke_color, position bottom middle
caption_add / caption_remove start, end, text or from_transcript: true, start?, end? / cue flattened: placement auto = mask over the detected caption band
caption_placement `mode band mask
title text, sub?, color?, band_color?, y?, height?, size? text: "" removes the band
theme `theme editorial mono
effect_add effect, start, end? / duration?, params {}, id? effect id / alias / zh label from the catalogue; params validated (unknown dropped, clamped, required checked)
effect_update id, start?, end?, shift?, duration?, params? move / retime / change params
effect_remove id
cover `t, text, style card plain
export_add / export_remove target (platform[:orientation], 3:4, 9:16, 16:9), `layout auto band`
reset everything back to the original (undoable)

3a. Transcript cuts (Descript-style editing)

Section titled “3a. Transcript cuts (Descript-style editing)”

Word indices refer to the cached transcript (transcript.json words [{w, t, te, p?}], paths.transcript in show); show.words_sig is its signature (sha1 of [w, t, te] per word, 10 hex; null while nothing is transcribed - show never transcribes).

op form meaning
{op: cut, words: [i0, i1], sig?, why: transcript} cut words i0..i1 (inclusive): start = W[i0].t, end = W[i1].te, then snap_cut (word-safe edges; if the audio refinement would drop a selected word the plain word times are used). sig != the current words_sig -> stale-words (nothing applied); indices not ints / out of range / reversed -> bad-param
{op: cut, gap: i, keep: 0.25, sig?, why: pause} shorten the pause between word i and i+1, keeping keep s (0.05-2, default 0.25) of it: cut [W[i].te + keep/2, W[i+1].t - keep/2], no word snapping; a pause shorter than keep + 0.05 -> bad-param
{op: cut, start, end, why: filler} a plain range (marks, AI) - snapped as before

Every cut edge (unless no_snap, or no audio) then moves to the nearest zero crossing of the audio the renderer cuts (the master in pipeline mode) within ±5 ms (outputs.zero_cross(samples, sr, t, t0) on ~20 ms decoded with one tiny ffmpeg call per edge; any failure leaves the edge): values[i].zero_cross: true. Cut times keep 5 decimals. Flattened + snap_captions: each edge of a transcript cut (words or why: transcript) moves to the nearest burned caption line boundary within 0.3 s (lines approximated by subs.cues_from_words(W, max_chars=14)); an edge that cannot snap keeps its place and adds warning cut-splits-burned-caption {at} (values[i].caption_snap true / false).

why is free text; the desk sends transcript, filler, pause. values[i] of a cut: {cut [a, b], words (what was said in it), snapped, kept_s, word_range?, gap?, gap_s?, keep_s?, zero_cross?, caption_snap?}; its history line is op-cut {start, end, why, said}.

One Apply = one step, captions and effects re-timed in it. When a step holds a cut with why transcript / filler / pause (or a words / gap cut), derived ops are computed against the final state of the step and appended to the SAME step (a single undo restores everything):

  • captions (pipeline cues): a cue with a word cut by this step loses the cut words (caption_text {cue, text: join_words(remaining), force: true}); nothing left (or no edited span) -> caption_remove; under 2 characters left -> removed and its text appended to the previous kept cue. Added captions are only re-mapped by the render.
  • effects (visual, overlapping a new cut): edited length < 0.4 s or fully cut -> effect_remove; pop-words whose word (its text among the words in its window, else the word at its start) was cut -> effect_remove; partly cut -> nothing (the render maps it), reported as trimmed. sfx-placement starting inside a new cut -> effect_remove. music-bed, joins, grade, end fade and the progress bar are never touched.
  • step.retimed (also on history.steps[].retimed): {captions {retimed, shortened, removed}, effects {trimmed [{id, effect, label {en, zh}, from_s, to_s}], removed [{id, effect, label}]}, sfx_removed, targets [primary, ...exports]} - every version updates: the cut happens before the canvas work of every target. Derived ops describe as op-caption-text {cue, text, force}, op-caption-remove {cue}, op-effect-remove {id}.

Marks (show.marks, from the cached transcript only, at most 500, never fails show): [{kind filler|pause|lowconf, i0, i1, text, save_s, group}] - fillers = cleanup.detect without audio, rows of kind filler not kept (group = the normalized filler, e.g. 那个; save_s = its length); pauses = gaps over 0.6 s between words i0 and i1 = i0 + 1 (text “1.2s”, save_s = gap - 0.25, group pause; cut it with {gap: i0}); lowconf = words with an ASR probability p < 0.5 (when the transcript has it).

preview-edl (output preview-edl --project P --output O --ops JSON --json, outputs.preview_edl(d, output, ops)): what the renderer would keep if the draft ops were applied; nothing is written (the transcript is read from the cache). Ops are normalized + folded in order exactly like edit (same snapping); a failing op is skipped and reported. -> {ok, output, keep [[a, b]] (outputs.segments, the renderer's own), cuts [[a, b]], duration (edited s), source_duration, ops [{index, op (normalized), describe, value}], dropped [{index, op, error}], warnings}. No derived caption / effect ops (they never change the kept ranges).

ai: output ai --instruction "..." [--apply] (or edit --op ai): the routed model (task output_edit of vstudio.llm; claude-code / codex / API / local routes all work) gets the output facts, caps, current state, the op list, the effect catalogue and the captions / transcript, and answers {ops, summary}. Every proposed op goes through the same validator as edit; invented effects / ops / params, out-of-range values and ops the caps forbid are dropped with their error. Without --apply nothing changes (the desk shows proposed for confirmation, then calls edit --ops with ops, or ai --apply). No model configured: a literal-phrase fallback (speed, head / tail trim, platform exports, LUFS, progress bar, fade) with warning no-model.

Selective revert (output revert --step <id>): cancels ONE earlier step and keeps every step after it. It is recorded as a new step {op: revert, step} (revert_of on the step; undo / redo it like any step; reverting a revert brings its target back); the state is folded without the cancelled step (history.steps[].reverted). Refused with revert-conflict {steps, n} when a later step builds on it (edits / removes an effect or added caption it created, removes a cut by index after it changed the cut list, or a later reset): the honest alternative is undo back to it. Also unknown-step, already-reverted.

Context (ai --context '{"range": [12, 18], "cues": ["c4"], "effect": "fx2"}'): what the creator points at (timeline selection, caption cues, one effect instance). Validated (bad-context, unknown-effect-instance), passed to the model as focus (cue text and the effect instance spelled out) with “this / here / 这段 mean the focus”; the literal fallback understands “剪掉这段 / cut this” and “zoom” on the range. Echoed as context.

Chat transcript (chat.json, output chat): every ai call records a turn (--no-record to skip) with its proposals, provider, model, cost and seconds, status draft; edit --turn ID marks it applied with the step id; revert of that step marks it reverted. The desk adds its own turns (slash-command cards) with chat --add JSON and patches status with chat --turn ID --set JSON. show returns chat so the conversation is rebuilt after a reopen.

Project-level AI (python -m vstudio.project ai --project P --instruction T [--outputs all|a,b] [--context JSON] [--timeout 120] [--json | --json-events], projai.py): one request for every output (“remove the series label from all clips”). (1) read every targeted output (cached transcripts only); (2) a cheap rule check BEFORE any model call: removing / replacing text burned into a flattened output, or restyling burned captions, can never be done by the editor, so those outputs get needs_rerender at once (no model, < 2 s) - evidence needed: a text noun (标题 / 水印 / label …), the text in the folder’s own scripts, or a cached transcript that does not say it (what is SAID goes to the model as a cut); text this editor added (title band, an effect’s text, an added caption) is removed with plain ops; (3) the rest goes to ONE model call (task output_edit) with every output’s summary, answered as groups per output, each op validated against that output’s caps. --timeout bounds each provider; a timed-out / failed provider falls back along the route’s chain (fallback event + fallback in the result). Nothing is applied: the desk applies a group with output edit (one undo step per output). needs_rerender = {outputs, titles, code burned-text | burned-captions-restyle | model-needs-rerender, reason, targets, paths [{kind rerender-scripts {files [{file, line, text}], scripts, prompt} | regenerate {items, command} | re-export, message, actions [{kind copy-prompt {prompt} | open-file {file, line} | regenerate {items} | reveal {file}, label}]}]}. --json-events: {event: stage, stage: read|check|ask|plan, n, provider?, elapsed}, {event: partial, needs_rerender, groups} (the rule answer, sent BEFORE the model call so the desk shows it at once), {event: fallback, from, to, code, error}, then {event: done, result} (or {event: failed, code, ...}, exit 5). A failed model call keeps the rule answer (warning llm-failed, failed {provider, code}).

4. Effects (output effects --json [--no-thumbs])

Section titled “4. Effects (output effects --json [--no-thumbs])”

Each id is a row of the effects registry (vstudio.effects); {id, label {en, zh}, description {en, zh}, kind, stage, params (JSON-Schema subset with x-zh), required, default_dur, aliases, thumbnail}; thumbnails are sample frames rendered once into $VSTUDIO_CACHE/output_fx/.

id zh stage params
pop-words 弹出大字 frame text, color, size, angle, anim, x, y
stacking-stamps 印章 frame text, angle, scale, anim, x, y
punch-in 推镜放大 frame scale, in_dur, out_dur, hold return
quote-card 金句卡 frame text, sub, speaker, width, theme, anim, x, y
callout-bubble 标注气泡 / 箭头 frame text, arrow_x, arrow_y, theme, anim, x, y
chapter-card 章节卡 frame title, index, total, accent
notes-panel 记笔记面板 frame title, bullets, theme, width, anim, x, y
overlay-images 贴纸 / 标签 frame image, text, style tag
badge 角标 frame text, color, scale, anim, x, y
red-box 框选高亮 frame x, y, w, h, color, width, anim
progress-bar-pil 进度条 frame style line
marker-sweep 马克笔划重点 frame text (【kw】), size, surface auto
chapter-rule 章节细线 frame label, title, index, width, anim, x, y
number-counter 数字滚动 frame value, label, prefix, suffix, count, size, anim, x, y
lower-third 人名条 frame name, role, scale, anim, x, y
sfx-placement 音效 audio name (the synthesized bank), gain
music-bed 背景音乐 audio file, duck_db, music_lufs (whole clip)
xfade-joins 转场 timeline transition (any vstudio.xfade name), duration: at a cut join = an xfade, elsewhere a flash / dip
end-fade 结尾淡出 timeline duration
vlog-grade 调色 timeline sat, contrast, warm

5. Render (output render [--quality preview|final] [--targets primary,douyin:vertical|all] [--json-events])

Section titled “5. Render (output render [--quality preview|final] [--targets primary,douyin:vertical|all] [--json-events])”

Stages per target: canvas (reframe the master / file to the target canvas) -> timeline (trim, cuts, joins, speed, grade, end fade) -> audio (SFX, music bed, two-pass loudness) -> final (frame pass: mask / band layout, punch-in, overlays, cards, title band, captions, progress bar; encode) -> cover. Each stage key hashes its input key + only the ops it uses, so a caption / overlay change re-runs final only, an SFX change audio + final, a cut everything from timeline; an unchanged render (or one undone back to) is all cache hits. preview = same resolution, low bitrate (ultrafast, CRF 30, 2 Mb/s); final = the platform delivery encode. show.renders says which renders are fresh for the current state. Every render writes .vstudio/status.json heartbeats (output-edit:<stage>, progress during the frame pass) into the edit folder and the owning project / work folder; the owner’s previous record is restored afterwards.

All commands take --json (paths absolute). Every user-facing message (caps notes, warnings, refusals, op descriptions) is {code, params, message (English), message_zh}: the desk localises by code + params; content (captions, titles, cover text) stays in the project’s content language.

command stdout
output list --project P `{ok, dir, kind project
output show --project P --output O {ok, output {id, file, mode, canvas, fps, duration, platform, master}, caps, caps_notes, state, captions [{id, start, end, text, original, edited, removed, added}], effects [{id, effect, start, end, params, label, edited [a, b]}], timeline {segments, joins, speed, duration}, history {steps [{id, at, by, note, describe, retimed}], undo, redo}, renders [{target, quality, file, fresh}], warnings, paths {doc, dir, renders, transcript}, words_sig, marks}
output edit ... --ops JSON the show document + {step {id, ops, describe, retimed?}, values [per op], warnings}
output preview-edl ... --ops JSON {ok, output, keep, cuts, duration, source_duration, ops, dropped, warnings} (nothing written)
output ai ... --instruction T [--context JSON] [--apply] {ok, context, proposed [{op, normalized, describe, why, warnings}], ops, dropped [{op, error}], summary, provider, model, cost_usd, seconds, warnings, turn, applied, step?}
output render ... {ok, output, mode, quality, targets [{target, file, cover, canvas, layout, duration, key, cached, stages [{stage, key, cached, seconds}], warnings}], seconds}; --json-events: target-start, stage-done, target-done, render-done lines; --with-ops JSON: a before / after preview of ops that are not applied, into renders/<target>.compare.mp4 (compare: true; edit.json and the manifest untouched)
output revert ... --step ID the show document + {step, reverted}
output chat ... [--add JSON / --turn ID --set JSON] {ok, turns} / {ok, turn}
`output undo redo …`
output effects {ok, effects [...], n}
ai --project P --instruction T [--outputs] `{ok, scope project, answer changes

Errors: exit 5 + {ok: false, error, code, params, message, message_zh}. Codes: unknown-owner, unknown-output, unknown-op, bad-op, no-ops, unknown-effect, unknown-effect-instance, duplicate-effect, bad-param, bad-time, too-short, cut-no-word, unknown-cut, no-audio, no-words, transcribe-failed, captions-not-ours, unknown-cue, empty-text, not-faithful, placement-pipeline, unknown-target, nothing-to-undo, nothing-to-redo, unknown-step, already-reverted, revert-conflict, bad-context, unknown-turn, llm-failed, llm-bad-json, render-failed, stale-words {expected, got}. Warnings / notes: flattened, captions-add-only, relayout-crops-burned, no-master, master-mismatch, source-changed, placement-auto, style-added-only, band-pipeline, effect-cut-away, param-adjusted, unknown-param, no-model, length, cut-splits-burned-caption {at}. Op descriptions: op-<op> (e.g. op-effect-add {effect, start, end}, op-revert {step, what}).