interview-qa
Cette page n’est pas encore traduite : voici la version anglaise.
Use when: someone hands over a conversation recording (podcast, interview, coaching call, office hours)
and wants Q&A clips: each clip opens on a question and keeps the answer tight (“把这期访谈切成问答”,
“一问一答的切片”, “cut this interview into Q&A clips with Chinese and English subtitles”). For free-form
highlights of a call, or gallery-view calls that need call-clips’ stacked / trio layouts and name-label
blur, use call-clips (this recipe writes its segment rows: work/segments.yaml).
Inputs → Outputs: talk.mp4 → per pair out/<qNN>/<platform>-<orientation>.mp4 (the question as a card
and / or the asker’s own voice, the answer trimmed, the question as a persistent header, role labels
“主持人 / 嘉宾” or “Host / Guest” at each speaker’s first line, bilingual captions, optional face mask),
a designed question cover per platform shape, <qNN>.<lang>.srt/.vtt, out/report.json (firstpass).
Defaults (format interview-qa)
Section titled “Defaults (format interview-qa)”“按「访谈 / 播客问答切片」默认:回答 1.1x,不加 hook,以问题开头,严格去口癖 / 气口,说话人标签用角色,中英双语字幕,
嘉宾露脸先问”. Never print a guessed name: speakers are roles unless she gives names (names: S1=…).
The consent checkpoint (never automatic) asks whether the guests agreed to publication.
export PYTHONPATH="$VSTUDIO/lib:$PYTHONPATH"python -m vstudio.project new talks/ep7 --recipe interview-qa --input source=talk.mp4 \ [--param question=both] [--param mask=sticker]python -m vstudio.project run --dir talks/ep7 # stops at "审核问答" and "嘉宾同意"By hand:
- Plan:
python -m vstudio.qa plan --source talk.mp4 --out work/pairs.json --segments work/segments.yaml [--speakers 2] [--diarize auto|pyannote|voice|text] [--names-csv S1=Ziyun] [--count 8]. Speakers: localpyannote.audiowhen installed with a local pipeline (VSTUDIO_DIARIZE_MODEL), else gallery tiles (--tiles host=x,y,w,h guest=…, mouth movement), else voice clustering, else text (questions = the asker). The asker with the most questions is the host. Pairs: a question run of the host, then the other speaker until the host takes the floor again, a long pause or--max. Answers lose an opening “great question / 好问题” and go throughvstudio.cleanup(tight):answer.windowsare the kept pieces. - Review
work/pairs.json(checkpointpairs): question text, windows, roles;"edited": true. - Render:
python -m vstudio.qa render work/pairs.json --source talk.mp4 --platforms xiaohongshu:full,douyin --question audio|card|both --subtitles bilingual [--mask sticker --mask-regions-csv "x,y,w,h"] [--speed 1.1] --out out. - Check
out/report.jsonand three frames: the question opening, a label, a masked face.
Captions, translation, glossary and SRT / VTT tracks work as in workflows/lesson-clips (vstudio.bilingual).
The question card stays on screen for its reading time (~3.5 words / 6 characters per second + 1 s, 2.5-6 s).
Voice clustering handles 2-3 clearly different voices; with similar voices check the roles in the review.