preproduction: script writing + pronunciation drills
Use when: before anything is recorded: drafting or revising a spoken script for a short (60–180 s) or
mid-length (3–10 min) video, fixing a flat hook or a generic ending, or preparing the narrator to say the hard
words in a locked script. Inputs: a topic / notes / draft. Outputs: SCRIPT.md (locked, linted) and an
optional pronunciation_drill.{m4a,wav,md} shadowing pack. Downstream: workflows/slides (visuals),
workflows/cover (the hook doubles as cover text), workflows/polish (final export).
Generic craft lives in references/script_craft.md. The creator’s own voice
(stance, preferred and banned phrases, insight signposts, pace, TTS voice) lives in persona voice.en.* (English
scripts) / voice.zh.* (中文口播) and tts.*; see the Voice schema below and
examples/persona.voice.example.yaml. Read both before writing.
Run from the project folder; $VSTUDIO = repo root.
A. Script
Section titled “A. Script”- Premise first (no prose yet). Answer in one or two sentences each: the one takeaway (≤15 words), what the viewer wrongly believes now, the mental tool they keep, why the creator finds it interesting. If the takeaway doesn’t fit in 15 words, stop and narrow the topic.
- Format + structure. Pick short / long-short / mid (craft §1). List sections with a one-line takeaway each; check that the takeaways tell a story in order.
- Draft in continuous spoken prose, in the persona voice for the script’s language (
voice.<lang>.persona,.phrases_prefer,.rules; see Voice schema). Overshoot length. Mark the one insight paragraph and signpost it (voice.<lang>.signposts). Write the hook last with the formula: you-statement → surprising claim → anchoring number. - Read-aloud pass (non-negotiable). Break tongue-trippers, connect disguised bullet lists, cut padding, doubled definitions, forward references.
- Lint:
The language is auto-detected (
Terminal window python3 $VSTUDIO/workflows/preproduction/scripts/lint_script.py SCRIPT.md --format shortpython3 $VSTUDIO/workflows/preproduction/scripts/lint_script.py SCRIPT.md --platform douyin # see Platforms--lang auto|en|zh; zh when CJK is ≥30% of the letters) and picks thevoice.<lang>profile and rule set. English errors: em-dashes, parentheses (if the persona bans them), trope openers, generic CTAs, authority framing, AI-tell words, meta/navigation lines, self-narration about the medium (“in this video”, “next slide”), filler reassurance (“don’t worry if…”), emoji,voice.en.phrases_avoid. Warnings: padding (“so what I want to say is”), sentences > 25 words, possible fragments, digits under 10, missing insight signpost, word count vs. target atvoice.en.wpm. 中文 errors: 破折号, 括号 (if banned), 口播套路开头 (家人们, 今天给大家分享, 话不多说…), 求赞求关注 CTA in the closing (点赞关注, 一键三连, 下期见…), authority framing (很多人不知道…), AI-tell words (赋能, 闭环, 值得注意的是…), meta lines (这期视频, 后面会讲…), filler reassurance (别担心, 听起来有点复杂…), emoji,voice.zh.phrases_avoid. Warnings: sentence > 40 字, missing signpost, 字数 vs. target atvoice.zh.cpm. The English-only heuristics (fragments, digits under 10, average words per sentence) are skipped for zh. Lint is a floor, not the judge: still run the self-check below. - Lock. Hand the hook sentence to the cover headline and the first slide. From here the script is frozen:
don’t rewrite it during recording (that is how a two-hour session becomes five). End-to-end recipe:
references/SOP_SHORT_VIDEO.md.
Time budgets per phase, the full pre-lock checklist and worked before/after revisions (hook ladder, paragraph,
closings A/B/C, 中文口播) are in references/script_craft.md §7, §8, §10.
SCRIPT.md format that other workflows read: ## SECTION headings, the spoken text as indented blocks,
optional **Delivery:** notes and ZH: translation lines (ignored by the linter). See
examples/script.example.md (synthetic).
Self-check before lock
Section titled “Self-check before lock”- Zero em-dashes; every sentence has a subject and a verb
- Hook ≤3 sentences (≤6 mid-length) with “you” and a concrete number or year; no throat-clearing
- Voice matches
voice.<lang>.persona; no banned openers, closers orvoice.<lang>.phrases_avoid - Not a consulted authority, no false modesty
- Exactly one signposted insight paragraph that reframes the topic at a deeper layer
- Closing uses pattern A/B/C/D; last line is screenshot-worthy alone
- No AI-tell words, meta/navigation lines or filler reassurance (lint clean)
- Word count fits the format; read aloud once with no stumbles
Voice schema
Section titled “Voice schema”One voice: block in persona.local.yaml, with a sub-profile per script language. lint_script.py resolves each
key as voice.<lang>.<key> → flat legacy voice.<key> → default, so a persona with only the flat keys lints exactly
as before (exception: the flat wpm is English words/min and is not inherited by zh). Full example:
examples/persona.voice.example.yaml.
Key (under voice.en / voice.zh; flat voice.* = legacy fallback) |
Default | Read by |
|---|---|---|
persona |
(none) | you / the LLM when drafting (stance: en = curious learner-educator, zh = confident practitioner) |
domain |
(none) | LLM-facing only |
language (flat legacy only) |
en |
doc only; the sub-block name now carries the language |
en.wpm (flat wpm) |
165 words/min | lint_script.py: seconds estimate, length target |
zh.cpm (or zh.wpm, read as 字/min) |
270 字/min (4.5 字/s) | lint_script.py: seconds estimate, length target |
rules |
[] | lint_script.py: “parenthes” / “括号” in a rule turns the parentheses check on; rest LLM-facing |
phrases_prefer |
[] | LLM-facing only |
phrases_avoid |
[] | lint_script.py: ERROR on each hit (on top of the built-in generic en / zh trope lists) |
signposts |
[] | lint_script.py: counts as the insight signpost (plus built-in generic ones) |
en.tts.<engine>_voice |
→ tts.<engine>_voice |
make_drill.py (drills are English) |
zh.tts.<engine>_voice |
(none) | documented for zh TTS previews; no preproduction tool reads it yet |
top-level tts.<engine>_voice (kokoro, edge, openai) |
af_heart / en-US-AriaNeural / cedar | vstudio.tts (make_drill.py, previews) |
Tests or one-off overrides: point VSTUDIO_PERSONA=/path/override.yaml at a YAML file; it is merged over the persona.
Platforms
Section titled “Platforms”lint_script.py --platform <name>[:orientation] (xiaohongshu, douyin, tiktok, youtube, youtube-shorts,
bilibili; aliases like xhs, 抖音, shorts work) turns the length check into a platform check:
- target range =
vstudio.platform.profile(name).length.sweet(seconds) × pace:voice.en.wpmwords/min for English,voice.zh.cpm字/min for Chinese. E.g. 抖音 sweet 15–60 s × 270 字/min → 68–270 字; YouTube 420–1200 s × 165 wpm → 1155–3300 words. - the estimated seconds go through
platform.check_length: over the hard max → ERROR; outside the sweet spot or under the minimum → warning. Override sweet/max per platform in personaplatforms.<name>.length. - precedence:
--platformwins;--formatis only used for length when no--platformis given (a note is printed if you pass both). Without either, the classicshorttarget applies, so existing runs are unchanged.
python3 $VSTUDIO/workflows/preproduction/scripts/lint_script.py SCRIPT.md --platform youtube-shortspython3 $VSTUDIO/workflows/preproduction/scripts/lint_script.py 口播.md --platform douyin # zh auto-detectedpython3 $VSTUDIO/workflows/preproduction/scripts/lint_script.py SCRIPT.md --platform xhs --lang enNothing else in preproduction depends on the platform (no canvas, captions or audio here); the rendered video’s
per-platform export happens downstream (python -m vstudio.export).
B. Pronunciation drill (after the script is locked)
Section titled “B. Pronunciation drill (after the script is locked)”For a narrator who wants to practise the hard words of this script before recording (non-native accent, tricky stress, names and jargon). Not a general pronunciation course: words that aren’t in the script add nothing, and more than ~12 words makes the drill too long; split it into two runs instead. The drill audio is a practice file only; it never goes into the video.
- Pick 5–10 words (8 is the sweet spot, the script refuses >12): read the script aloud, mark every stumble or unclear stress, prioritise by frequency in the script > visibility > pitfall risk. Every word must be in the locked script; drills map to what is about to be recorded, not general vocabulary.
- Write the list as JSON (
word, sentence, ipa, pitfall, bad, good) or TXT (word | sentence | ipa | pitfall):cp $VSTUDIO/workflows/preproduction/examples/drill_words.example.json work/drill_words.json - Generate:
Per word: slow word (0.65×) → 1.2 s → slow word → 1.2 s → script sentence (0.85×) → 2.2 s; spoken intro first. Output
Terminal window python3 $VSTUDIO/workflows/preproduction/scripts/make_drill.py work/drill_words.json \-o work/pronunciation_drill --script SCRIPT.mdpronunciation_drill.wavand.m4a(both two-pass loudnorm topersona.audio.loudness_lufs, 48 kHz stereo; the m4a is AAC 96k, AirPods-ready), and.mdreference card (IPA, pitfall, BAD/GOOD, source sentence; words missing from the script are flagged).--dry-runwrites only the card. Length ≈ 6 + n × 9.6 s (real runs land a bit longer with natural pacing). - TTS engine (
--engine autopicks the first available):kokoro— Kokoro-82M viamlx-audioon Apple Silicon (pip install mlx-audio), local, default voiceaf_heart.edge—pip install edge-tts; any OS, free, needs network; defaulten-US-AriaNeural.openai—OPENAI_API_KEY; any OS;gpt-4o-mini-ttswithspeed(--model tts-1for the old model); default voicecedar. All engines go throughvstudio.tts.synth: clips are cached in$VSTUDIO_CACHE/tts/, so re-running a drill after editing one sentence only synthesises that sentence. Voice per engine:--voice, else personavoice.en.tts.<engine>_voice(drills are English), elsetts.<engine>_voice, else the engine default. Use the same voice as any TTS preview so the reference stays consistent.
Drill self-check
Section titled “Drill self-check”- Every word appears in the locked script (
--scriptflags the ones that don’t) - 5–10 words; each has IPA, a one-line pitfall and the source sentence (BAD/GOOD when the mistake is typical)
- Audio loudness-normalised (the script does this; check the printed LUFS)
-
.m4awell under ~1 MB per minute (larger usually means a wrong sample rate or bitrate) - The
.mdcard sits next to the audio; practise 2–3 times before recording
- mlx-audio Kokoro can crash on one specific sentence at one speed (
broadcast_shapes ... cannot be broadcast, surfaced as “mlx_audio produced no wav”). Reword the sentence slightly or nudge--normal(0.85 → 0.86). - The craft is written for English scripts. For 中文口播 keep the structural rules (hook formula, one insight,
closing patterns, no em-dashes); the linter switches to character counts (
voice.zh.cpm, default 270 字/min ≈ 4.5 字/s) and a 40 字 sentence limit. - Don’t run the drill on a draft: words change, and the narrator ends up drilling sentences they won’t say.