Ir al contenido

interview-qa

Esta página aún no está traducida; aquí tienes la versión en inglés.

Use when: someone hands over a conversation recording (podcast, interview, coaching call, office hours) and wants Q&A clips: each clip opens on a question and keeps the answer tight (“把这期访谈切成问答”, “一问一答的切片”, “cut this interview into Q&A clips with Chinese and English subtitles”). For free-form highlights of a call, or gallery-view calls that need call-clips’ stacked / trio layouts and name-label blur, use call-clips (this recipe writes its segment rows: work/segments.yaml).

Inputs → Outputs: talk.mp4 → per pair out/<qNN>/<platform>-<orientation>.mp4 (the question as a card and / or the asker’s own voice, the answer trimmed, the question as a persistent header, role labels “主持人 / 嘉宾” or “Host / Guest” at each speaker’s first line, bilingual captions, optional face mask), a designed question cover per platform shape, <qNN>.<lang>.srt/.vtt, out/report.json (firstpass).

“按「访谈 / 播客问答切片」默认:回答 1.1x,不加 hook,以问题开头,严格去口癖 / 气口,说话人标签用角色,中英双语字幕, 嘉宾露脸先问”. Never print a guessed name: speakers are roles unless she gives names (names: S1=…). The consent checkpoint (never automatic) asks whether the guests agreed to publication.

Terminal window
export PYTHONPATH="$VSTUDIO/lib:$PYTHONPATH"
python -m vstudio.project new talks/ep7 --recipe interview-qa --input source=talk.mp4 \
[--param question=both] [--param mask=sticker]
python -m vstudio.project run --dir talks/ep7 # stops at "审核问答" and "嘉宾同意"

By hand:

  1. Plan: python -m vstudio.qa plan --source talk.mp4 --out work/pairs.json --segments work/segments.yaml [--speakers 2] [--diarize auto|pyannote|voice|text] [--names-csv S1=Ziyun] [--count 8]. Speakers: local pyannote.audio when installed with a local pipeline (VSTUDIO_DIARIZE_MODEL), else gallery tiles (--tiles host=x,y,w,h guest=…, mouth movement), else voice clustering, else text (questions = the asker). The asker with the most questions is the host. Pairs: a question run of the host, then the other speaker until the host takes the floor again, a long pause or --max. Answers lose an opening “great question / 好问题” and go through vstudio.cleanup (tight): answer.windows are the kept pieces.
  2. Review work/pairs.json (checkpoint pairs): question text, windows, roles; "edited": true.
  3. Render: python -m vstudio.qa render work/pairs.json --source talk.mp4 --platforms xiaohongshu:full,douyin --question audio|card|both --subtitles bilingual [--mask sticker --mask-regions-csv "x,y,w,h"] [--speed 1.1] --out out.
  4. Check out/report.json and three frames: the question opening, a label, a masked face.

Captions, translation, glossary and SRT / VTT tracks work as in workflows/lesson-clips (vstudio.bilingual). The question card stays on screen for its reading time (~3.5 words / 6 characters per second + 1 s, 2.5-6 s). Voice clustering handles 2-3 clearly different voices; with similar voices check the roles in the review.