cover: cover images for short and mid-length videos
Use when: a video is cut and needs a cover / thumbnail that replaces a dead first frame and doubles as the
platform thumbnail (小红书, Shorts, TikTok, Reels, B站, YouTube). Three patterns: A 4-frame diagonal collage of
stills, B the speaker’s matted face over a 2×2 grid of slides, C premium split cover (retouched face photo
left, dark panel right with quote / title / thumbnail / chips) for 4:3 and 16:9 posts.
Inputs: the cut video (and/or a face photo, slide PNGs). Outputs: cover.png (1080×1920, A/B) or
cover-4x3.jpg + cover-16x9.jpg (C), plus one file per platform on request (see Platforms). Feed the cover to
workflows/polish (--cover) or python -m vstudio.export --cover.
Run from the project folder; $VSTUDIO = repo root. Working files go in work/.
Prerequisites
Section titled “Prerequisites”./install.shdone (fonts in the cache,selfie_segmenter+face_landmarkermodels).- Chrome / Chromium / Edge for HTML renders (found via
$CHROME, PATH, or standard install paths; a broken binary is skipped) — orpip install playwright && playwright install chromiumas fallback. - Optional, Pattern B only:
torchif you want the RVM matting engine (GPL-3.0, downloaded at runtime; see below).
Pick a pattern
Section titled “Pick a pattern”| Pattern | Use when | Bias |
|---|---|---|
| A collage | preview content variety; face not needed on the thumbnail | info density |
| B face on quadrants | personal brand; the face drives CTR | personal connection |
| C split cover | horizontal post (小红书 4:3 / YouTube 16:9), quote or “亲测” angle, one hero thumbnail | premium, editorial |
A new channel usually gets more from B (the face builds recognition); an established one can lean on A
(content density). Not for: static image posts (use an image editor), or a long-form YouTube thumbnail that should
follow YouTube’s own conventions (big face, 3–5 words); --platform youtube only resizes these layouts to 1280x720.
Pattern A — collage
Section titled “Pattern A — collage”- Find 4 moments (1 hook, 2–3 content, optionally 1 face):
python3 $VSTUDIO/workflows/cover/scripts/extract_frames.py sheet my-talk.mp4 work/cover_src --every 5→work/cover_src/contact_sheet.jpgwith timestamps (--save-framesalso writes full-sizecand_<t>.png). Talking-head source?extract_frames.py pick my-talk.mp4 work/cover_src --top 6ranks frames by smile / eyes open / centred face (vstudio.cover.score_frames) →pick<N>_<t>.png+picks_sheet.jpg. - Pre-crop the cells (crop boxes are fractions of the decoded frame; append
:faceto use the face box):python3 $VSTUDIO/workflows/cover/scripts/extract_frames.py collage my-talk.mp4 work/cover_src 2.5 37 76 80:faceDefaults fit a “slide on top, face below” layout; change--slide-box/--face-boxfor other layouts. cp $VSTUDIO/workflows/cover/templates/cover_collage.template.html work/cover.html, edit the tag, badges, headline and pills.python3 $VSTUDIO/workflows/cover/scripts/render_cover.py work/cover.html -o work/cover.png(1080x1920; add--platform xiaohongshu --platform douyin ...for per-platform sizes, see Platforms)
Pattern B — face on quadrants
Section titled “Pattern B — face on quadrants”- Pick a talking frame (mid-word, mouth slightly open, eyes on camera; idle frames look posed) from the contact sheet
(or the
pickranking above). - Crop the face region, stopping above any burned-in caption band:
python3 $VSTUDIO/workflows/cover/scripts/extract_frames.py face my-talk.mp4 work/cover_src/face_src.png --t 70 --box 0,0.5,1,0.948 - Optional retouch:
python3 -m vstudio.retouch work/cover_src/face_src.png work/cover_src/face_src.png --slim .05 --eye .04(--preset none|natural|daily|glamfor makeup,--faces allfor group shots; seereferences/RETOUCH.md) (PYTHONPATH=$VSTUDIO/libif vstudio isn’t pip-installed). - Matte:
python3 $VSTUDIO/workflows/cover/scripts/matte.py work/cover_src/face_src.png work/cover_face.png --trim --preview- default engine: MediaPipe selfie segmenter + guided-filter edge refinement (Apache-2.0, fast, CPU).
--engine rvm: Robust Video Matting, cleaner hair. GPL-3.0, not vendored:torch.hubdownloads it at runtime under its own licence. Needs torch; uses MPS/CUDA when available; warm-up passes are built in (RVM is recurrent, a single cold pass gives a soft alpha).- Caption still visible at the bottom?
--crop-bottom 0.1. - Check
work/cover_face.check.png(alpha over a checkerboard) for halos.
- Copy 4 clean slides (from
workflows/slides) towork/cover_src/bg_tl.png … bg_br.png. cp $VSTUDIO/workflows/cover/templates/cover_face_quadrants.template.html work/cover.html; set.face-wrapwidth/height to the same aspect as the matted PNG (matte.py prints it). Otherwiseobject-fit:containshrinks the face and leaves empty space. Example: PNG 1080×770 (1.40) → 980×700.- Render as in A (
render_cover.py).
Pattern C — premium split cover (retouched face left, dark panel right)
Section titled “Pattern C — premium split cover (retouched face left, dark panel right)”- Grab the best face frame at full resolution (any aspect; it is scaled to the cover height) and retouch it:
python3 -m vstudio.retouch work/cover_src/face.png work/cover_src/face_retouched.png --slim .05 --eye .04 --makeup .5 --preset natural(or put"retouch": {...}in the config and split_cover.py does it, cached as*.retouched.png). - Optional hero thumbnail: a frame of the thing the video shows (an explainer frame, a product shot).
- Write
work/split_cover.jsonfrom$VSTUDIO/workflows/cover/examples/split_cover.example.json:quote(serif italic, highlight colour) +byline — a short third-party line that frames the post.title.lines(CJK bold, auto-shrinks to the panel) withtitle.highlightsubstrings inbrand.highlight.thumbnail.path+cropfractions (rounded corners, drop shadow, hairline outline).chips(outlined pills,ink|dim|highlight|accent),stamp(rotatedbrand.accenttag over the photo, e.g. 亲测),corner_tag(bottom-rightbrand.highlighttag, e.g. 记笔记 ↓).outputs: one entry per size, e.g. 1440×1080 withphoto_w760 (4:3) and 1920×1080 withphoto_w900 (16:9).face_x: "auto"centres the photo crop on the detected face (vstudio.face); a number (0–1) overrides.
python3 $VSTUDIO/workflows/cover/scripts/split_cover.py work/split_cover.json [--platforms xiaohongshu:horizontal,xiaohongshu,douyin,youtube]- Look at both sizes. 小红书 shows a centre crop of horizontal covers in the feed, so keep the face and the first
title line inside the middle ~75 %. Title length rules:
persona.platforms.<platform>.title_max.
Platforms
Section titled “Platforms”Cover sizes come from vstudio.platform.cover_size (PLATFORMS.md):
| platform | cover | feed shows |
|---|---|---|
小红书 (xiaohongshu, 3:4 / 9:16 posts) |
1080x1440 | whole cover |
小红书 horizontal (xiaohongshu:horizontal) |
1920x1080 | centre 4:3 (x 240..1680) |
抖音 / TikTok (douyin, tiktok) |
1080x1920 | profile grid: centre 3:4 |
| YouTube Shorts | 1080x1920 | whole cover |
| YouTube | 1280x720, ≤ 2 MB; keep titles left of x 1100 (timestamp) | whole cover |
| B站 | 1146x717 (16:10) | whole cover |
- A / B (templates):
render_cover.py page.html -o cover.png --platform xiaohongshu --platform douyin --platform youtubelays the page out at a design size of the same aspect (short side 1080), scales type with--fs, then resizes to the exact cover size;<out>.<platform>-<orientation>.pngeach, +.feed.jpgwhere the feed crops. Platform files are always suffixed, even for a single--platform(alsodouyin, which is 1080x1920 too), so they never overwrite the plain-ocover; render once without--platformif you need the unsuffixedcover.png. - C (split cover): an output with
"platform": "..."takes its size from the profile. Portrait sizes (3:4, 9:16) use the stacked layout (photo band on top, panel below). For 小红书 16:9 the design is laid out feed-safe: the whole 4:3 design sits in the centre crop and the photo runs on to the left edge ("feed_safe": falseturns it off); every platform output with a feed crop also writes<path>.feed.jpg.--platforms a,b,cadds outputs. - V-track talking-head covers:
workflows/talkinghead/scripts/vertical/cover.py config.py --platform ...(same sizes).
Rules (all patterns)
Section titled “Rules (all patterns)”- Exact canvas sizes: 1080×1920 for A/B by default (per platform with
render_cover.py --platform); whateveroutputssays for C. - Accent. The collage / face-quadrant templates use
--cover-accent= personacover.accent(default teal#2dd4bf, their original look;render_cover.py --accent '#ff2442'for one render). They no longer followbrand.accent, which stays the on-video red. Pattern C (split cover) still usesbrand.accent/brand.highlight(stamp, chips, title highlight) so it matches the video’s graphics. Plainpython -m vstudio.renderon a template renders teal at 1080x1920. - Headline pattern for A/B:
Subject <span class="punch">is/isn't [contrarian punch].</span>(accent italic punch). It should be the same sentence as the script hook (seeworkflows/preproduction). - Fonts: the templates use
@font-faceonassets/fonts/*whichvstudio.rendercopies from the repo font cache. Never point at system fonts. - Each quadrant / cell must show distinct content; four near-identical frames read as a glitch.
- Trust the decoded frame size over ffprobe metadata (it has reported 720×1280 for a 1080×1920 file).
- The cover only replaces the first ~1 s of picture (
workflows/polish --cover); audio is untouched.
Self-check
Section titled “Self-check”- Canvas size exact; PNG opens with no broken transparency
- Accent/highlight colours match the slides and on-video graphics
- B: matte has no halo and no caption strip;
.face-wrapaspect = PNG aspect - C: face and title line 1 survive a centre crop; quote fits on one line; chips don’t hit the right edge
- Top tag / headline don’t cover the focal content
- Headline = the script’s hook sentence, punch in accent italic (A/B)
- File size sane (a 1080×1920 PNG lands around 0.5–1 MB; YouTube needs ≤ 2 MB)
- After upload: the platform actually shows the custom cover (some swap back to a video frame)