Skip to content

Sound: beats, SFX vocabulary, grammar, levels

How to make cuts land on the music and how to place sound effects so a fast edit feels deliberate instead of busy. Code: lib/vstudio/beats.py (analysis, grids, cut plans, energy arcs) and the SFX section of lib/vstudio/audio.py (synthesised bank, place_sfx, cue sheets).

Every sound in the SFX bank is synthesised from sine and noise maths, so there are no audio files to license.

  1. Music first, if there is music. Analyse it before planning shots: b = beats.analyze("bgm.mp3"). Save the analysis (b.save("work/beats.json")): it is the audit trail for every cut.
  2. Picture next. Plan sections with beats.energy_arc(duration, kind, beats=b), then cut times with b.cut_plan(...). Write shot boundaries as beat numbers (b.beat_t(n)), not frame numbers.
  3. Sound effects last, once the picture is locked. Any change to shot lengths or order means the cue sheet has to be rebuilt. Rebuilding is cheap when cues come from cue_sheet_for(...) with the same event list.
  4. Verify after rendering: beats.verify(cut_times, b, tol_frames=3, fps=30).

Keep times as float seconds all the way. Round to frames only at the last step, because rounding early lets errors pile up.

Field Meaning
bpm, period, offset Least-squares straight-line fit t = offset + n * period through the tracked beats. The tracker’s own tempo number can be off by a few percent. The fitted line is not.
residual_ms, residual_p90_ms, grid_ok Fit quality. The grid is accepted when p90 <= 15 ms and max <= 30 ms, after dropping up to 10 % stray beats. Otherwise beats keeps the raw tracked beats, which happens with tempo drift, live drummers or DJ mixes.
tempo_check 1/2x, 1x and 2x candidate grids, each scored by how much kick energy lands on it. A grid that sits on the off-beats is shifted by half a beat.
onsets Per-band transients: kick < 150 Hz, snare 150-2500 Hz, hat > 5 kHz, as (t, strength) pairs.
downbeats, downbeat_phase Bar = 4 beats. The phase is the beat position with the strongest kicks.
rms, sections Loudness curve (50 ms hop). Section boundaries come from spectral novelty and are snapped to bars, each with energy 0-1 and a level.
hits The top-N strongest transients, at least 2 beats apart, weighted by loudness.

Backends: librosa’s tracker when it is importable, otherwise a pure-numpy tracker: spectral-flux onset envelope, autocorrelation tempo with a 120 BPM prior, then a dynamic-programming beat path. Both backends share everything after tracking: attack refinement on a 2 ms curve, grid fit, tempo check, bands and sections.

Grid or transient?

  • For dense, regular cuts (every beat, every 2 beats), use the grid: b.beat_t(n) and b.cut_plan(...).
  • For sparse accents, such as a freeze-frame or the one big slam, use the real transient from b.hits or b.onsets. A tiny grid drift gets exposed on an isolated accent.
  • The last frame of the video should land on the last real hit, with the RMS curve confirming the music has stopped. Don’t end on a reverb tail.
  • Strong accents are almost always on whole beats. Put a slam on a half beat only if the onset data shows a hit there.

Cut patterns (b.cut_plan(n, start, end, pattern)). every-beat, every-2 and every-bar are steady. accelerate uses shrinking gaps of 2, 1.5, 1, 0.75, 0.5, 0.375 and 0.25 beats, which is 32/24/16/12/8/6/4 frames at 112 BPM and 30 fps, and resolves on a bar line. drop holds two bar cuts, then cuts on every beat from the drop.

Verify. verify reports two errors per cut:

  • Audio error: the designed seconds against the beat. This checks the analysis.
  • Frame error: the frame-rounded time against the beat. This checks the frame rate.

Up to 3 frames of frame error passes, and 1.5 frames or less is ideal. 30 fps cannot be more precise than about 17 ms. If every cut is off by the same amount, look for an output audio offset first (AAC encoder priming is about 1024-2112 samples). Measure it once per render pipeline by cross-correlating a sharp SFX against the rendered audio. Keep that offset as its own constant and never fold it into the analysed offset.

3. Vocabulary (synthesised, 48 kHz, audio.sfx_bank())

Section titled “3. Vocabulary (synthesised, 48 kHz, audio.sfx_bank())”
Sound Use it for Peak in sample Peak level at base gain
soft (alias transition) scene / place change, one per change 0.47 s -14 dBFS
whoosh, whoosh_b camera move, push, pan 0.15-0.18 s -19 dBFS
swoosh, swoosh_b short: title or card entrance, small push, zoom 0.13 s -11 dBFS
whip (alias whip_pan) whip pan; peaks at the cut, stops dead 0.24 s -10 dBFS
impact, impact_b (alias boom) landing, big slam, logo stamp; the loudest moment start -7 dBFS
stamp (alias thud) small landing: sticker, badge, stamp start -5 dBFS
riser build into a finale or reveal, 2.4 s end (2.2 s) -12 dBFS
sparkle tail after an impact, glow, “ta-da” 0.03 s -12 dBFS
shutter, shutter_b (alias camera) photo moment, freeze-frame snapshot start -9 dBFS
pop, pop_b list items, stickers appearing start -16 dBFS
tick, tick_b counters, map pins, small steps start -15 dBFS
ding check mark, “done” start -19 dBFS
typewriter, typewriter_b one key / one character start -12 dBFS
typing typing reveal: trim with dur to the text animation start -13 dBFS
scratch (alias record_scratch) comedic “wait, what?” freeze 0.28 s -11 dBFS
stop (alias tapestop) everything stops, slow-down gag start -10 dBFS

The *_b sounds are alternation partners. audio.write_sfx("assets/sfx") writes every sound (aliases skipped) as stereo wav for HyperFrames <audio> clips.

Choose by genre, not by event. A cinematic promo uses whoosh, impact, riser, sparkle and transition, and skips cartoon sounds. In cue_sheet_for(kind="promo"), pop becomes tick and ding becomes sparkle. A fun travel vlog can use pop, scratch, shutter and ding. Test each sound in the finished cut, not on its own. Ask yourself: with eyes closed, does it sound like this kind of video, or like a mobile game?

Real actions get matching sounds. Typing on screen gets keys, a photo gets a shutter, a sticker landing gets a stamp. A generic whoosh won’t cover a distinctive action. Trim long sounds to the length of the action (dur).

4. Grammar (audio.cue_sheet_for(cuts, reveals, kind))

Section titled “4. Grammar (audio.cue_sheet_for(cuts, reveals, kind))”

A cue sheet is a plain list: [{t, sfx, gain_db, dur, note}]. t is where the sound’s peak lands, and note says which on-screen action it belongs to. No cue goes in without a reason.

Rule What the code does
One soft transition per scene change A scene cut gets soft. Transition-type cues closer than 0.3 s collapse into the most important one, so a cut-derived cue beats a reveal-derived one.
Beat cuts are carried by the music beat, montage, jump and hard cuts get no SFX. The drums already mark them.
Whoosh on camera moves move/push get whoosh, zoom gets swoosh, whip gets whip.
Impact on landings landing gets impact, snapped to the nearest beat within 0.12 s when beats= is given.
Finale = riser -> impact -> sparkle finale at t: the riser peaks at t (-2 dB), the impact hits at t, and the sparkle follows 0.6 s later (-3 dB). This is the one sentence never to break.
Repeated hits The same sound repeated within 1.5 s alternates with its _b variant and steps down 1.5 dB per repeat (max -9 dB). The run should sound countable, not like a machine gun. If the hits get too dense to count, replace them with one swoosh.
Density cap At most N cues in any 2 s window: travel-fun and promo 3, talking-head and story 2. A finale triple counts once. The most important cues stay: impact/riser > sparkle/stop/whip > transitions > pops/ticks.
Kind profiles CUE_PROFILES: the cap, an overall gain offset (talking-head -3 dB, story -2 dB) and vocabulary swaps.

Then x = audio.render_cue_sheet(cues, total) gives an (n, 2) float32 stem, peak-aligned. Write it with audio.write_wav and mix it with the voice and music.

Many sounds don’t peak at their first sample: a riser peaks at its end, a whip just before its cut-off, a soft transition halfway through. If you place them by their start, every hit sounds late. place_sfx with dict events, and render_cue_sheet, shift each sound so that its 5 ms-RMS peak (audio.sfx_peak) lands on t. The default is peak_align=True for dict events and False for legacy (t, name) tuples. For external samples, measure the peak once with sfx_peak and subtract it the same way.

Layer Level
Voice -16 LUFS stem (persona audio.voice_lufs)
Music under voice -30 LUFS bed, ducked a further ~10 dB while the voice speaks (mix_bed)
Music-only montage Bring the bed up to about -18 to -20 LUFS. The music is the voice here.
Regular SFX Peaks around -18 to -10 dBFS: under the voice, level with the drums
Accent SFX The finale impact is the loudest SFX (about -7 dBFS peak), used 2-3 times per video at most
Final Two-pass loudnorm of the whole mix to -14 LUFS, true peak -1.5 dBTP (loudnorm_2pass)

True peak and AAC: an encoded output (.mp4 / .m4a) is measured after the encode and corrected. A .wav output is usually muxed to AAC later, which adds ~0.2-0.4 dB of peak, so loudnorm_2pass(..., "mix.wav") limits WAV_HEADROOM_DB (0.5 dB) under the target (headroom=0 for an exact lossless deliverable). After a custom mux, audio.ensure_loudness(final.mp4) re-measures and re-normalises only if it missed.

Use loudness for importance. The most important beat gets the loudest SFX, and repeated small sounds sit lowest. If the music already hits hard, let the drums do most of the work and give SFX only to actions that exist only on screen. Gain is a multiplier, not a target: a quiet external sample needs normalising (or a different sample) before it can sit at these levels. Check the level in the rendered file.

Deliver two versions from the same timeline: with music and without music. The SFX and the voice stay in both. Platforms mute or swap licensed tracks, and a no-music version can be re-scored without a re-edit.