Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,15 @@

## Unreleased

**`bin/vh pace`: fixes from the review of #72**
- Why: an independent review of #72 (merged before its findings were in) found that an ASR boundary off by more than about 0.1 s could drop syllables or move one behind a pause, a crash when a line has no audio, a missing key found only after the take was written, and some smaller gaps. A second independent review, of this PR, found that the breath rule had no loudness test, that a line with no audio added silence and got zero-width times, that swapped spans repeated audio, and that CI could not catch a revert of the word-time alignment; those are fixed here too.
- Line edges: a line's voice is everything voiced between the silences that part it from its neighbours, each found within 0.2 s of the ASR boundary (at least 0.12 s long when there is one; shorter dips are between syllables); the ASR start and end only locate those silences. Before, on the synthetic take, starts 0.2 s late or ends 0.3 s early lost syllables in every trial, and a boundary 0.1 s inside a syllable moved it behind a 0.39 s pause; now nothing is lost in 40 trials each (a start or an end off by up to 0.5 s either way), and a boundary 0.05–0.2 s inside a syllable leaves it where it was. A whole boundary shifted 0.4 s or more one way can still put a few syllables in the neighbouring line (nothing is lost). A blip at a line's edge under 0.12 s, set off by 0.06 s of quiet and 15 dB under the line's speech (one per side: a breath), stays in the audio but is not counted as speech in the pace or the pauses; a line opening with three short loud syllables measures 5.17 against a true 5.07 (6.36 with a length-only rule). Timelines whose spans go back in time (a hand edit) are refused: cutting at their silences repeated up to 3 s of speech.
- Word times follow the audio: `atempo` stretches a line unevenly (pauses and speech by different amounts), so the loudness envelopes before and after are aligned (dynamic time warping within ±0.15 s of the even stretch). Mapped word starts on the synthetic take are within 8 ms of the real onsets (48 kHz; 27 ms at 22.05 kHz in a line slowed to ×0.92), against 40–60 ms with the average stretch; a line's first word starts and its last word ends where the speech does. On the 2-minute take, mapped line starts sit on the voice onsets (median 0, max 0.03 s).
- A line nothing was heard for (missing from the audio, or under the voice threshold, like a whisper) is named and left alone: its stretch of the take is copied as it is, unpaced, and the line gets a 0.1 s span with its words spread across it. Before, a skipped last line crashed, a skipped middle line got zero-width times and an extra clamped pause, and a whispered line was replaced by silence. A "……" line keeps the pause it stands for, with that stretch's own audio. The analysis hop is a whole number of samples at every rate (at 22.05 kHz it drifted 0.23%), and the alignment keeps only its band (a 120 s line took 1.2 GB, now 0.23 GB).
- `--align gemini` without `GEMINI_API_KEY`, or `--align whisper` without mlx-whisper, stops before anything is written (it wrote the take and skipped the captions). `--restore --dry-run` is refused (it restored). `--ref` on a line without words is an error instead of a ZeroDivisionError. The language may follow the flags. A missing `script.txt`, or line ids not in it, is said (those lines count as one block, and their pause tags are not known). A failed first transcription exits with a message instead of a traceback. The `--restore` refusal no longer suggests deleting the `.raw` files, which may be the only copy of the original.
- A take without word times is transcribed once and kept in `audio/pace.<lang>.asr.json` under its sha256, so the `.raw` timeline stays byte-identical (it was rewritten with the transcription); `--dry-run` transcribes but writes nothing.
- tools/ci.sh: the synthetic take is 22.05 kHz with a "……" line, and mapped words must be within 35 ms of their syllables (27 ms now, 61 ms without the alignment). New cases: ASR starts 0.2 s late and ends 0.3 s early, a line break 0.1 s inside a syllable, the last line and a middle line with no audio, swapped spans, a quiet breath and three short loud syllables, word times only in the cache (and none at all), no `script.txt`, `--target` / `--extra` / the language after the flags, `--restore --dry-run`, `--align gemini` without a key, `--ref` on a line without words, and `voiceover.wav` / `timeline.json` restored too. Run against #72's `pace.py`, and against this PR's first commit, they fail where the two reviews said.

**`bin/vh pace`: bring the pace of the lines in one narration take closer, without synthesizing again**
- Why: a one-request Gemini take ran audibly fast and slow from line to line, and a project script evened it out on the take itself (the steps went into playbook/04 in #70). The maintainer asked for a command, and for it to stay gentle ("别太严格这个语速"): it is optional, moves each line part of the way, leaves the lines that are already close, and its numbers are a reference for the ear.
- Per line: lines split at the silence nearest the ASR boundary (an aligner had put one line's end and the next one's start at the same instant, inside the first one's last syllable), then cut points refined from the waveform (an ASR start placed late by a misheard first syllable no longer clips it); pauses inside a line longer than `--pause` 0.26 s shortened to it, unless the script asked for them (`<short pause>`, `<long pause>`, `<#s#>`); pace = spoken units per voiced second (CJK characters with digits as they are read, or English syllables); tempo `(reference / pace) ^ --strength` (0.8) within `--clamp` (0.92–1.18), through ffmpeg `atempo`; lines within `--tolerance` (5%) of the reference, or under 4 units or 0.5 s, keep their pace. The reference is the median of the lines, `--ref id` or `--ref first-last` (a passage a person approved), or `--target N`.
Expand Down
2 changes: 1 addition & 1 deletion playbook/04-audio.md
Original file line number Diff line number Diff line change
Expand Up @@ -219,7 +219,7 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词
**语速的依据和用法**:表里是句内语速,不算句间停顿。普通话正常朗读、去掉 150 ms 以上的停顿后约 5.2 字/秒:七种语言的朗读语料里,汉语是 5.18 音节/秒,一个汉字一个音节(Pellegrino, Coupé & Marsico 2011,*Language* 87(3),[doi:10.1353/lan.2011.0057](https://doi.org/10.1353/lan.2011.0057))。电话闲聊按总时长算是每分钟 228–247 字,约 3.8–4.1 字/秒(Yuan, Liberman & Cieri 2006,Interspeech,[doi:10.21437/interspeech.2006-204](https://doi.org/10.21437/interspeech.2006-204))。知识科普的 4.5–5.5 围着正常朗读,原理讲解慢一档。倒推稿子字数时乘 0.85,留出句间停顿【综合】:45 s 的知识短视频约 170–210 字,3 分钟的讲解约 540–690 字。TTS 的实际语速看音色:本节第 4 小节比较 `duck_ratio` 用的那段测试旁白(0.25 s 间距,本地 Qwen3-TTS 0.6B 的 Serena,7 句 107 字)句内约 4.1 字/秒。所以字数只用来起稿,时长以 `timeline.json` 的实测为准(硬规则 2)。

**一条 take 里快慢差得明显时,可以逐句拉近**:整稿一次合成、不写逐句指示,语气能统一,句与句的语速却可能差不少。句子之间本来就该有快慢,人听着不别扭就不用动;听得出忽快忽慢时,不用重合成,用 `bin/vh pace` 在这条 take 上逐句处理。先加 `--dry-run` 看每句的语速和要拉的倍数,再写:
1. 起止按波形定:从对齐给的句首往回走,走到声音静下来为止(ASR 给的句首不一定准,第一个字听错时常常偏晚,照着剪会切掉它),句尾同样往后走;
1. 起止按波形定:ASR 给的边界只用来找两句之间的那段静音,一句的人声就是两段静音之间有声的部分(ASR 给的句首句尾不一定准,第一个字听错时常常偏晚,照着剪会切掉它);
2. 句内过长的停顿压到约 0.25 s,切口两侧各淡几毫秒;稿子里写了停顿标签的句子不压;
3. 每句用 `atempo`(不变调)往同一个语速拉近一些,不必拉齐。参照默认取各句的中位数;人听过、认可了某一段(比如开头),用 `--ref` 指定它。只校正一部分:倍数取(参照 ÷ 这句的语速)的 0.8 次方,限在 0.92–1.18,差得不多的句子不动,自然的起伏留着;
4. 句间停顿用 take 自己的,只把太短、太长的按句尾标点和段落收进一个范围,个别地方用 `--extra` 加减;拼接别用 `anullsrc` 加 `concat`(见上文"旁白晚一点进来")。写回 `voiceover.<lang>.wav` 和 `timeline.<lang>.json`,逐字时间跟着换算过去,原件留成 `.raw`,字幕跟着重做。
Expand Down
Loading
Loading