photo-story
Esta página aún no está traducida; aquí tienes la versión en inglés.
Use when: Turn your photos / short clips plus a narration script (or just a music track: MODE = "music", cuts on bars) into an effect-rich story video (museum visit, travel, art history, product story) with bilingual EN/中文 subtitles, a running header with section progress, chapter cards, a cover and post copy, at any canvas size (3:4, 9:16, 16:9) or laid out for a platform (PLATFORM). Narration by OpenAI TTS or your own cloned voice. Pure Python (PIL + OpenCV + ffmpeg), driven by one spec.py.
Use when someone has a folder of photos (and maybe a few phone clips) and a story to tell over them: “make a narrated video of my museum visit / trip / this artist / this product”, with Ken-Burns moves, collages, maps, timelines, quotes, highlight circles, film looks, and burned-in bilingual subtitles. Not for talking-head footage (see talkinghead workflows) or HTML/HyperFrames motion graphics.
Inputs → Outputs: my_photos/*.jpg, optional my_clips/*.mov, optional music bed, a spec.py
(titles, sections, script lines with EN+中文 cues, shot list) → out/story.mp4 (H.264 + AAC, bt709-tagged, two-pass loudnorm),
out/cover.png, out/subtitles.srt, out/transcript.md, out/voice.mp3, out/post.md.
Scripts live in $VSTUDIO/workflows/photo-story/scripts/photostory/; run them from the project
folder. Paths inside the spec are relative to the spec file. Needs ffmpeg, numpy, opencv-python,
Pillow, soundfile, OPENAI_API_KEY for TTS (no openai package needed), and a whisper backend (mlx_whisper on Apple Silicon, else
faster_whisper, else the OpenAI whisper-1 API). Fonts come from vstudio.config.font (./install.sh).
中文 titles, chapter names and quotes use the cjk-serif role (Noto Serif SC from install.sh). For another serif
(e.g. a system Songti) set persona fonts.cjk-serif: /path/Songti.ttc#1 (a .ttc#N picks face N; index 0 of
Songti is the Black face). If no serif is found the titles fall back to cjk-bold (sans) and the render prints a
!! photo-story: no 'cjk-serif' font warning - fix it before delivering a 文艺 piece.
Pipeline
Section titled “Pipeline”- Gather media. Put photos in
my_photos/(names are the shot ids:photo07.jpg→"photo07"), clips inmy_clips/(IMG_1234.MOVcan be referenced as"v1234"). Still frames from a clip: extract them tomy_photos/with ffmpeg and use them as photos. - Write the spec. Copy
examples/demo_spec.py→work/spec.py. FillTITLE_*,SECTIONS(_EN),CANVAS, thenSCRIPT: one item per narration clip,(section, [(en, zh), ...], [(src, weight, motion, opts), ...]). Write**word**for highlights. Keep each EN cue ≤ ~90 chars and 中文 ≤ ~38 chars. Put content data in spec tables, not code:ROUTE(cities/legs),TIMELINE(range/ticks),FOCUS,VCROP,VEQ,FILM_SECTIONS. - Layout check before paying for TTS. Without
timing.json, durations are estimated (silent):Stills go toTerminal window python3 $VSTUDIO/workflows/photo-story/scripts/photostory/render.py work/spec.py --stills 2,10,25,40work/stills/at 1/3 size (--fullfor full size). Look at each one. Check mid-transition times too. - Voice (skip in
MODE = "music"). OpenAI:OPENAI_API_KEYmust be in the environment. Your own voice:VOICE = dict(engine="clone")(see “Voice clone” below).Each unit is taken up to 3× and the take whose transcript best matches the script wins (log showsTerminal window python3 $VSTUDIO/workflows/photo-story/scripts/photostory/tts.py work/spec.py --sample # pick a voicepython3 $VSTUDIO/workflows/photo-story/scripts/photostory/tts.py work/spec.py # all units -> work/tts/python3 $VSTUDIO/workflows/photo-story/scripts/photostory/tts.py work/spec.py 4 9 # redo units 4 and 9score). The chosen take is cached inwork/tts/cache/by text+voice+model+instructions, so editing shots never re-bills; editing a line re-synthesises only that line. Synthesis goes throughvstudio.tts(48 kHz mono wav), transcription throughvstudio.asr. - Preview, then render.
Rendering is single-process CPU work. The demo previewed 3 s at 1080x1440 in about 4 s on Apple Silicon. 1620x2160 and heavy shots (tilt, loupe, ink) are slower, so start long renders in the background.
Terminal window python3 .../render.py work/spec.py --preview 20 --from 60 # -> work/cache/preview.mp4python3 .../render.py work/spec.py # -> spec OUT - Cover + side outputs.
Terminal window python3 .../cover.py work/spec.py # COVER dict -> out/cover.pngpython3 .../export.py work/spec.py # subtitles.srt, transcript.md, voice.mp3 (-16 LUFS), post.md
- narration (default): the TTS
timing.jsondrives the timeline; optionalBGMbed under the voice. - music (
MODE = "music"orrender.py --mode music, no TTS):beats.analyze(BGM)drives it. Every shot gets a whole number of grid units by weight, so every cut lands on a bar (or beat). Withlengthset, each shot gets at least its kind’s minimum time (MIN_SECONDS: collage / grid / rows / deck 2.5 s, film / route / quote 3 s, split 2 s, others 1 unit;MUSIC min_seconds={"collage": 3}or a shot’smin_s=overrides). Sections are sized by weight but never below their shots’ minimums, so a late chapter is not squeezed into 1-bar shots; alengthtoo short for the minimums is raised and the log prints! length: ... set MUSIC length=N- follow it or drop a shot. Withoutlength(weight xperunits, as before) short multi-picture shots are only warned about (raise their weight). Spec sections start on the music’s own section changes when one is within a quarter section (chapter cards + header follow). The(en, zh)lines become quiet title text (中文 serif + English italic, 0.6 s fades) spread over each item’s shots;captions=Falsehides them,"subs"uses the speech-subtitle style. Lines may be empty:(0, [], shots).MODE = "music"; BGM = "music/bed.mp3"MUSIC = dict(grid="bar", per=1.0, length=None, start=0.0, tr="fade", trd=None, sections=True,captions="title", lufs=-16, backend="auto", min_seconds=None)grid:bar(calm, 文艺片) ·beat·2(every 2 beats) · any number of beats.per= units per weight;length= target seconds instead. Default transition = a fade of half a grid unit (0.25-1.4 s) starting on the downbeat; per-shottr=/trd=still win. Render printsbeats.verifyand writescache/music_cuts.json+cache/beats.json(audit trail). The music is normalised as the main track (lufs, then the final loudnorm), fades out over the last bar. Ambient / rubato tracks have no steady grid (grid rawin the log): cuts then follow the tracked beats.
Video clip audio
Section titled “Video clip audio”Video shots are silent by default. ("vIMG_1234", 1, "still", dict(audio="keep", gain=-3)) keeps the clip’s own sound
in the foreground (the music dips by CLIP_DUCK, default -10 dB); audio="duck" keeps it as ambience (set to
AMBIENT_LUFS -26, dipping by AMBIENT_DUCK -12 dB while the narration speaks); audio="mute" = default. gain is dB
on top. Kept clips are levelled to CLIP_LUFS (persona audio.voice_lufs, -16). Clip sound follows off=/speed= and
fades with the picture transitions. iPhone HLG/PQ clips are converted to SDR first (avconvert on macOS, else ffmpeg tone-map).
Clips with speech: 气口 / filler / repeat cleanup (opt-in). The shared tool vstudio.cleanup
(references/CLEANUP.md) is used, never a local cutter: dict(audio="keep", cleanup="gentle") on the shot, or
CLEANUP = "gentle" in the spec for every audio="keep" clip (ambient duck / mute clips are left alone; a shot’s
cleanup=False opts out). Values: True (= gentle) / "gentle" / "standard" / "tight" / "pauses" (气口 only) /
dict(profile=, reply="确认 3,5 / 保留 7", approve=, keep=, all_confirm=). It runs analyze --ranges off-<clip end>
then apply (auto edits only unless the creator approves more), and picture + sound both read the cleaned clip
from 0. EDL + review sheet: <CACHE>/cleanup/<clip>_<off>.cleanup.json / .review.md. Transcript: words= (file next
to the clip), <clip>.words.json, else ASR.
Voice clone (your own voice)
Section titled “Voice clone (your own voice)”Offline Qwen3-TTS (Base) via mlx-audio on Apple Silicon, cloned from a 5-15 s recording of your voice:
# persona.local.yaml (git-ignored) - never commit the recording or put it inside the repotts: clone: ref_wav: ~/voice/ref.wav # clean speech, no music, one speaker ref_text: "the exact words spoken in ref.wav" model: mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16 # optional (default); ~4 GB, downloaded on first useSpec: VOICE = dict(engine="clone") (optional ref_wav=, ref_text=, model=, tries= 4, good= 0.93;
VOICE is merged over TTS), or tts.py spec.py --engine clone. Each unit is generated with seeds 1000*unit+try,
levelled to -20 dBFS RMS, transcribed by whisper and scored against the script; a speech rate < 1.6 or > 4.2 words/s
costs 0.2 (stuck / garbled takes). The best take wins; it is cached per unit by reference-audio hash + transcript +
model + text, and each seeded take is cached by vstudio.tts. Library use: tts.synth(text, engine="clone", seed=1).
Platforms
Section titled “Platforms”PLATFORM = "xiaohongshu:vertical" in the spec, or --platform <name[:orientation]> on render.py, cover.py,
export.py (douyin, tiktok, youtube, youtube-shorts, bilibili:horizontal, xiaohongshu:full|horizontal …).
From vstudio.platform: the canvas (overrides CANVAS), the header content starts at the safe-zone top and stays inside
the safe left/right margins, subtitles are fitted (size from the profile’s caption range, shrunk until both languages fit)
into caption_box (clear of the lower-right button column), the picture box sits between them (portrait) or under the
overlaid subtitles (landscape), labels stay above the subtitles and inside the safe margins, the final mix is loudnormed
to the profile’s LUFS / true peak, cover.py uses cover_size (e.g. YouTube 1280x720), post.md uses the platform’s
format, and render/export print check_length warnings. Without PLATFORM the old CANVAS layout is used unchanged
(length checks use persona platforms.default in the canvas orientation). The area below the caption box (the platform’s
description UI) stays ground-coloured.
Because the header, picture box and labels are laid out per canvas, the best multi-platform route is re-rendering per
platform (render.py spec.py --platform douyin --out out/douyin.mp4; cached shots make it cheap). For an extra
encode/cover/post pass, or a same-aspect variant, write a caption-free master and let vstudio.export burn captions:
python3 .../render.py work/spec.py --clean-master --out out/master.mp4 # + out/master.cues.jsonpython3 -m vstudio.export out/master.mp4 --platforms douyin,tiktok --cues out/master.cues.json --cover out/cover.pngexport.py also writes cues.json (text = 中文, alt = EN) next to the SRT.
Spec reference (all optional except SCRIPT)
Section titled “Spec reference (all optional except SCRIPT)”| key | meaning |
|---|---|
PLATFORM |
vstudio.platform profile, e.g. "xiaohongshu:vertical" (sets canvas + safe layout, LUFS, cover size) |
MODE, MUSIC |
"narration" (default) or "music" + its options (see Modes) |
VOICE |
TTS engine options merged over TTS; engine="clone" = your own voice |
CLIP_DUCK, AMBIENT_DUCK, CLIP_LUFS, AMBIENT_LUFS |
clip-audio levels / ducking (see Video clip audio) |
CLEANUP |
气口 / filler cleanup profile for every audio="keep" clip (default off; see Video clip audio) |
CANVAS |
"3:4" 1080x1440 · "3:4-hd" 1620x2160 · "9:16" 1080x1920 · "16:9" 1920x1080 · "WxH" |
LAYOUT |
header, sub (fraction of H or px), overlay_subs. Portrait default: header 13.9 %, picture box, sub band 16.7 %. Landscape default: header 15 % and subtitles overlaid on the picture over a gradient |
PALETTE |
accent ground paper ink mark route sub_en sub_zh (hex or RGB) |
TITLE_ZH/EN, SECTIONS, SECTIONS_EN |
running header + chapter cards (cards for sections ≥ 1, CHAPTER_CARD seconds, 0 = off) |
FILM_SECTIONS |
sections with the heavy memory-film look (grain, warm desat, flicker, scratches, dust) |
MEDIA |
images, videos folders and extensions |
FOCUS |
{img: ((cx,cy), zoom)} framing used inside collage/film/grid/split/rows/deck |
VCROP / VEQ / VEQ_FILTER |
per-clip crop (cx, cy, width_frac) / clips that get the grade filter. Crop fractions are of the displayed frame (rotation metadata applied: an iPhone portrait clip stored landscape with a -90 deg rotation is cropped as portrait, as in any player / contact sheet). Crop and grade are part of the clip cache key. |
ROUTE |
title, note, cities={name: dict(lon, lat, zh, year, side)}, legs={key: [(a,b),...]} |
TIMELINE |
start, end, ticks, labels, format for the tl= bar |
FILM_CAPTION |
caption under film-strip frames, {n} / {name} |
TTS |
dir voice model instructions tries good use_say sample_voices sample_line whisper faster_model language |
SAY |
spoken forms for numbers/names (alignment; also TTS input when use_say) |
BGM, BGM_LUFS, BGM_DUCK, OUT, AUDIO_BITRATE, FPS, PACING |
music bed (your own licensed file) at BGM_LUFS loudness (default persona audio.music_lufs, -30), BGM_DUCK dB while the voice speaks (default 0 = static bed), output; PACING = gap sec_gap tail lead end_fade. BGM_VOLUME is no longer used |
COVER, POST, VOCAB |
cover.py layout data / post copy (title intro outro tags platform use_persona_tags tag_set; platform picks the vstudio.publish format, default persona platforms.default; tags are the post’s own, persona publish.tags are appended unless use_persona_tags=False; tag_set="art" appends persona publish.tag_sets.art instead) / optional study table |
Shot sources
Section titled “Shot sources”"img" photo · "v<clip>" / "video:<clip>" (opts off, speed, grade, audio=keep|duck|mute, gain, cleanup, words) · collage:a,b,c · film:a,b,c ·
split:a|b (labels=) · grid:a,b,c,d · tilt:a (yaw=(from,to)) · deck:a,b,c · quote:a (q=(en, zh, who)) ·
route:<leg> or route:A>B>C · medal:a (ring="TEXT · ") · rows:a,b,c.
Photo motions: in out panL panR up down still flip; options c=(cx,cy), z=zoom, label="...", tr=<transition>, trd=seconds.
Overlays (any shot)
Section titled “Overlays (any shot)”fx=("sketch","develop","shimmer","dust","prick"), hl=(cx,cy,rx,ry) red-pen circle (hl_t),
tri=((x,y),(x,y),(x,y)) composition triangle (tri_label), loupe=[(x,y),...] travelling magnifier (loupe_mag),
tl=year or tl=(a,b) timeline bar, count=(n,"EN","中") count-up. Coordinates are 0-1 of the picture box.
Transitions (tr=)
Section titled “Transitions (tr=)”fade push whip flash zoom iris leak ink blinds tear slideup cut.
Gotchas
Section titled “Gotchas”- Transitions overlap the previous shot: each shot keeps rendering for the next shot’s transition time. Very short shots (< 1 s) with long transitions (leak 0.7 s) look busy.
- Weights, not seconds: shots share their script line’s slot (voice + gap) by weight. To hold a shot longer, give it more weight or split the line.
prickruns edge detection on the first frame it sees. Use it on astillshot.sketchlooks best on detailed photos withinorstill.flipmirrors the image half-way, which suits “drawn in mirror image” moments.- Labels sit bottom-left of the picture. In landscape with overlaid subtitles they move above the subtitle band.
- Chinese text in any serif/italic slot automatically switches to a CJK font, because STIX has no CJK glyphs.
- Highlights (
**x**) survive line wrapping. Lines are balanced so the last line is never a lone word. - Route maps are schematic: an equal-aspect projection of the given lon/lat, fitted into the box. Use
side="left"to keep labels from colliding. - Music mode with a click-free ambient track:
beats.analyzemay find no steady grid. Check the log (grid raw, p90 residual) andcache/music_cuts.json; pick a track with a pulse, orMUSIC.backend="numpy"|"librosa"to compare. - Landscape: the medal ring is shrunk and centred above the overlaid subtitles; the header puts the title left and the section progress right.
- Editing a line’s cue count invalidates that unit’s timing. Render falls back to an estimate for that unit and prints a warning. Re-run
tts.py <i>.