diff --git a/AGENTS.md b/AGENTS.md index e3a7ed3..aa491f8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -108,7 +108,7 @@ OpenVideoHarness/ ├── CLAUDE.md / AGENTS.md 入口(本文件;AGENTS.md 由 bin/vh sync-agents 生成,内容相同) ├── README.md / README.zh-CN.md 给人看的说明;CONTRIBUTING.md 是改本仓库时的规则;install.sh 一键安装 ├── LOCAL.md 本机环境(不入库;模板是 LOCAL.example.md) -├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk +├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · pace · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk ├── tools/ bin/vh 背后的脚本;audio/README.md 是声音命令的参数手册,desk/README.md 是审阅台的读取约定;ci.sh 是仓库自检 ├── skills/open-video-harness/ 轻量 skill:在任何目录把做视频的请求引到本仓库 ├── video-types/ 9 类视频(09 实验中):工作流、审美、禁止项、Prompt 增量块、自查重点、案例 diff --git a/CHANGELOG.md b/CHANGELOG.md index 80a994e..f21b763 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,15 @@ ## Unreleased +**`bin/vh pace`: bring the pace of the lines in one narration take closer, without synthesizing again** +- Why: a one-request Gemini take ran audibly fast and slow from line to line, and a project script evened it out on the take itself (the steps went into playbook/04 in #70). The maintainer asked for a command, and for it to stay gentle ("别太严格这个语速"): it is optional, moves each line part of the way, leaves the lines that are already close, and its numbers are a reference for the ear. +- Per line: lines split at the silence nearest the ASR boundary (an aligner had put one line's end and the next one's start at the same instant, inside the first one's last syllable), then cut points refined from the waveform (an ASR start placed late by a misheard first syllable no longer clips it); pauses inside a line longer than `--pause` 0.26 s shortened to it, unless the script asked for them (``, ``, `<#s#>`); pace = spoken units per voiced second (CJK characters with digits as they are read, or English syllables); tempo `(reference / pace) ^ --strength` (0.8) within `--clamp` (0.92–1.18), through ffmpeg `atempo`; lines within `--tolerance` (5%) of the reference, or under 4 units or 0.5 s, keep their pace. The reference is the median of the lines, `--ref id` or `--ref first-last` (a passage a person approved), or `--target N`. +- Pauses between lines stay the take's own, voice to voice, kept inside a range set by how the line ends (comma or none 0.12–0.40 s, full stop 0.25–0.70 s, end of a block 0.45–1.20 s); `--extra id=s` adds or takes away after a line. Joined in numpy, not with `anullsrc` + `concat`. +- Writes `voiceover..wav` and `timeline..json` back (and the `timeline.json` / `voiceover.wav` copies of that language), keeps the originals as `.raw`, and redoes the captions. A second run starts again from the `.raw` take, so settings never compound; a new `bin/vh tts` take replaces it (told apart by a sha256 in the timeline's `pace` block); `--restore` puts the original back; `--dry-run` only measures. A timeline on a beat grid and a dialogue are refused. +- Word times: `--align map` (default) carries every word through the cuts and the tempo change: offline, deterministic. `--align gemini` (one call per ≤ 9 min of audio) or `--align whisper` (a local mlx-whisper, Apple Silicon) transcribe the new take and place the lines with tts.py's `align_lines`; the same flag transcribes a take that has no word times yet. A failed transcription (a spent daily quota) keeps the paced take with the mapped times and exits 1. +- Measured: a synthetic take of 42 syllables with known onsets at 4.1–7.3 per second lost none, and the mapped word starts were 2 ms (median) and 40 ms (max) from the real onsets. A 2-minute, 27-line Gemini take went from 5.18–7.38 chars/s (sd 0.50) to 5.88–6.83 (sd 0.21) with 13 lines' tempo untouched, 129.5 s → 121.4 s; every mapped line start was within 0.04 s of the voice onset, a Gemini re-transcription agreed with them within 0.03 s (median), whisper's line starts came 0.19 s early (median); `bin/vh qa` counted 53 click warnings after against 57 before. +- playbook/04's paragraph names the command, tools/audio/README describes it, CLAUDE.md lists it. tools/ci.sh: a smoke test on a synthetic take (tone-burst syllables at different rates, word times known): nothing written by `--dry-run`, every syllable kept and every mapped word on its syllable, the spread of paces smaller, tempos inside the clamp, an inner pause capped and an asked-for one kept, pauses inside their ranges, two lines with one shared ASR boundary split at the silence, the originals kept, a second run byte-identical, `--restore` byte-identical to the original, a beat-grid timeline refused. + **`bin/vh tts --align gemini`: a spent daily quota fails at once instead of looking hung** - Why: Gemini's 429 for a spent per-day quota can ask for a short `retryDelay` (27 s, 11.8 s in reports on Google's forum), and tts.py took a 429 for the daily quota only when it asked for more than 90 s. Below that it waited and retried up to 5 times a call, 4 calls at a time, so a 2026-10 film's `--align gemini` looked hung (documented in the entry below). - tts.py reads the quotas a 429 names: when a `google.rpc.QuotaFailure` violation's `quotaId` or `quotaMetric` has `PerDay` in it (`GenerateRequestsPerDayPerProjectPerModel`, also when it is listed with the per-minute quota), the call fails at once, whatever delay it asks for, with a message that names the quota and says `--resume` continues the run once it resets. The take is kept as before. Per-minute 429s are waited out and retried as before, and a 429 that names no daily quota still fails at once only when it asks for more than 90 s. The same applies to `bin/vh voices`. diff --git a/CLAUDE.md b/CLAUDE.md index 8283e62..57a367b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -105,7 +105,7 @@ OpenVideoHarness/ ├── CLAUDE.md / AGENTS.md 入口(本文件;AGENTS.md 由 bin/vh sync-agents 生成,内容相同) ├── README.md / README.zh-CN.md 给人看的说明;CONTRIBUTING.md 是改本仓库时的规则;install.sh 一键安装 ├── LOCAL.md 本机环境(不入库;模板是 LOCAL.example.md) -├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk +├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · pace · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk ├── tools/ bin/vh 背后的脚本;audio/README.md 是声音命令的参数手册,desk/README.md 是审阅台的读取约定;ci.sh 是仓库自检 ├── skills/open-video-harness/ 轻量 skill:在任何目录把做视频的请求引到本仓库 ├── video-types/ 9 类视频(09 实验中):工作流、审美、禁止项、Prompt 增量块、自查重点、案例 diff --git a/bin/vh b/bin/vh index 8aea89f..682f456 100755 --- a/bin/vh +++ b/bin/vh @@ -33,6 +33,9 @@ # --align gemini: word timestamps + ASR-vs-script check per line; --join: several lines per request; # --resume: continue the last run after a failed synthesis or transcription, same script and command) # providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite | minimax +# pace [zh|en] [--ref [-] | --target N] [--strength 0.8] [--clamp 0.92,1.18] [--tolerance 0.05] [--pause 0.26] [--extra id=s,…] [--align map|gemini|whisper] [--dry-run] [--restore] +# optional: bring the pace of the lines in one take closer (part of the way, natural variation kept), no new synthesis; +# --dry-run only measures. The original stays as voiceover..raw.wav; captions are redone # voices list [lang] | design "" [--out f.wav] | delete Gemini voice library / voice design # captions [zh|en] [zh-max] [--media-end S] SRT (zh / en / bilingual) + captions.json from the TTS timeline; a caption under 1.8 s is # lengthened into the gap after it, up to the picture's end (--media-end, else index.html's data-duration, media/final.mp4, the timeline) @@ -680,6 +683,11 @@ cmd_tts() { uv run -q --no-project $deps python "$ROOT/tools/audio/tts.py" "$proj" --provider "$prov" ${lang:+--lang "$lang"} ${voice:+--voice "$voice"} ${instruct:+--instruct "$instruct"} "$@" } cmd_voices() { [ $# -gt 0 ] || set -- list; python3 "$ROOT/tools/audio/tts.py" voices "$@"; } +cmd_pace() { # numpy for the waveform; mlx-whisper only with --align whisper (Apple Silicon) + local a prev="" deps="--with numpy" + for a in "$@"; do case "$prev:$a" in --align:whisper|*:--align=whisper) deps="$deps --with mlx-whisper" ;; esac; prev=$a; done + uv run -q --no-project $deps python "$ROOT/tools/audio/pace.py" "$@" +} cmd_captions() { # --media-end S: where a caption too short for the 1.8 s floor may be lengthened to (default: the picture's length) local u="usage: bin/vh captions [zh|en] [zh-max] [--media-end S]" local proj=${1:?$u}; shift @@ -767,12 +775,13 @@ cmd_usage() { # one command's lines from the header above, with their continuati case "${2:-}" in -h|--help) # `bin/vh -h` prints its usage instead of running it with -h as a file, project or preset name case "$1" in captions|readcheck|review|storyboard|rhythm|cover-preview|voices|mix|recipes|sfx|qa) ;; # these print their own, fuller help tts) cmd_usage tts; echo; exec python3 "$ROOT/tools/audio/tts.py" -h ;; # + every flag, without building the mlx-audio environment first + pace) cmd_usage pace; echo; exec python3 "$ROOT/tools/audio/pace.py" -h ;; *) cmd_usage "$1" && exit 0 ;; esac ;; esac case "${1:-help}" in doctor) cmd_doctor ;; setup) cmd_setup ;; types) cmd_types ;; - tts) shift; cmd_tts "$@" ;; voices) shift; cmd_voices "$@" ;; captions) shift; cmd_captions "$@" ;; music) shift; cmd_music "$@" ;; sfx) shift; cmd_sfx "$@" ;; beats) shift; cmd_beats "$@" ;; mix) shift; cmd_mix "$@" ;; qa) shift; cmd_qa "$@" ;; mux) shift; cmd_mux "$@" ;; + tts) shift; cmd_tts "$@" ;; pace) shift; cmd_pace "$@" ;; voices) shift; cmd_voices "$@" ;; captions) shift; cmd_captions "$@" ;; music) shift; cmd_music "$@" ;; sfx) shift; cmd_sfx "$@" ;; beats) shift; cmd_beats "$@" ;; mix) shift; cmd_mix "$@" ;; qa) shift; cmd_qa "$@" ;; mux) shift; cmd_mux "$@" ;; install-skill) shift; cmd_install_skill "$@" ;; recipes) shift; python3 "$ROOT/tools/recipes.py" "$@" ;; review) shift; cmd_review "$@" ;; diff --git a/playbook/04-audio.md b/playbook/04-audio.md index 9a64627..fa803ab 100644 --- a/playbook/04-audio.md +++ b/playbook/04-audio.md @@ -218,13 +218,13 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词 **语速的依据和用法**:表里是句内语速,不算句间停顿。普通话正常朗读、去掉 150 ms 以上的停顿后约 5.2 字/秒:七种语言的朗读语料里,汉语是 5.18 音节/秒,一个汉字一个音节(Pellegrino, Coupé & Marsico 2011,*Language* 87(3),[doi:10.1353/lan.2011.0057](https://doi.org/10.1353/lan.2011.0057))。电话闲聊按总时长算是每分钟 228–247 字,约 3.8–4.1 字/秒(Yuan, Liberman & Cieri 2006,Interspeech,[doi:10.21437/interspeech.2006-204](https://doi.org/10.21437/interspeech.2006-204))。知识科普的 4.5–5.5 围着正常朗读,原理讲解慢一档。倒推稿子字数时乘 0.85,留出句间停顿【综合】:45 s 的知识短视频约 170–210 字,3 分钟的讲解约 540–690 字。TTS 的实际语速看音色:本节第 4 小节比较 `duck_ratio` 用的那段测试旁白(0.25 s 间距,本地 Qwen3-TTS 0.6B 的 Serena,7 句 107 字)句内约 4.1 字/秒。所以字数只用来起稿,时长以 `timeline.json` 的实测为准(硬规则 2)。 -**一条 take 里快慢差得明显时,可以逐句拉近**:整稿一次合成、不写逐句指示,语气能统一,句与句的语速却可能差不少。句子之间本来就该有快慢,人听着不别扭就不用动;听得出忽快忽慢时,不用重合成,在这条 take 上逐句处理(还没做成命令): +**一条 take 里快慢差得明显时,可以逐句拉近**:整稿一次合成、不写逐句指示,语气能统一,句与句的语速却可能差不少。句子之间本来就该有快慢,人听着不别扭就不用动;听得出忽快忽慢时,不用重合成,用 `bin/vh pace` 在这条 take 上逐句处理。先加 `--dry-run` 看每句的语速和要拉的倍数,再写: 1. 起止按波形定:从对齐给的句首往回走,走到声音静下来为止(ASR 给的句首不一定准,第一个字听错时常常偏晚,照着剪会切掉它),句尾同样往后走; -2. 句内过长的停顿压到约 0.25 s,切口两侧各淡几毫秒; -3. 每句用 `atempo`(不变调)往同一个语速拉近一些,不必拉齐;参照取人听过、认可的那一段(比如开头)。只校正一部分,例如倍数取(参照 ÷ 这句的语速)的 0.8 次方,限在 0.92–1.18,自然的起伏留着; -4. 按标点和段落重排句间停顿(拼接别用 `anullsrc` 加 `concat`,见上文"旁白晚一点进来"),写回 `voiceover..wav`;整条再对齐一次,更新 `timeline..json` 的逐字时间,再重跑 `bin/vh captions`。 +2. 句内过长的停顿压到约 0.25 s,切口两侧各淡几毫秒;稿子里写了停顿标签的句子不压; +3. 每句用 `atempo`(不变调)往同一个语速拉近一些,不必拉齐。参照默认取各句的中位数;人听过、认可了某一段(比如开头),用 `--ref` 指定它。只校正一部分:倍数取(参照 ÷ 这句的语速)的 0.8 次方,限在 0.92–1.18,差得不多的句子不动,自然的起伏留着; +4. 句间停顿用 take 自己的,只把太短、太长的按句尾标点和段落收进一个范围,个别地方用 `--extra` 加减;拼接别用 `anullsrc` 加 `concat`(见上文"旁白晚一点进来")。写回 `voiceover..wav` 和 `timeline..json`,逐字时间跟着换算过去,原件留成 `.raw`,字幕跟着重做。 -一支 2026-10 的片子,一条 Gemini 整稿各句原来 4.7–6.7 字/秒(第 1、2 步之后量),这样处理后在 5.6–6.1 之间,听起来还是同一条 take。数字只是参考,以人听着顺为准。 +一支 2026-10 的片子,一条 Gemini 整稿各句原来 4.7–6.7 字/秒(第 1、2 步之后量),手工这样处理后在 5.6–6.1 之间,听起来还是同一条 take;`bin/vh pace` 的默认值拉得比这轻。数字只是参考,以人听着顺为准。参数、实测和 `--restore` 见 `tools/audio/README.md` 的"逐句拉近语速"。 **要做 A/B**:`studio` 档位下,同一段旁白至少试两种导演方向,把两版都给人听。实测对比见本节开头:同样三句话,只给一个"纪录片旁白"的整体语气时听起来平;逐句导演、再卡上拍之后,才有起伏和节奏。 diff --git a/tools/audio/README.md b/tools/audio/README.md index 27dbf29..bcd259d 100644 --- a/tools/audio/README.md +++ b/tools/audio/README.md @@ -1,8 +1,8 @@ # 声音命令的参数手册 -`playbook/04-audio.md` 讲声音怎么做、怎么判断;这里是 `bin/vh` 各条声音命令的细节:参数、文件格式、实测数字和已知局限。用到哪一节查哪一节,不用通读。各步的实现另写在对应脚本的开头(`tts.py`、`captions.py`、`music.py`、`instruments.py`、`motifs.py`、`sfx.py`、`mix.py`、`qa.py`)。 +`playbook/04-audio.md` 讲声音怎么做、怎么判断;这里是 `bin/vh` 各条声音命令的细节:参数、文件格式、实测数字和已知局限。用到哪一节查哪一节,不用通读。各步的实现另写在对应脚本的开头(`tts.py`、`pace.py`、`captions.py`、`music.py`、`instruments.py`、`motifs.py`、`sfx.py`、`mix.py`、`qa.py`)。 -## 配音(`tts`、`captions`) +## 配音(`tts`、`pace`、`captions`) ### 句首句尾的静音 @@ -80,6 +80,28 @@ 所以动效类、要卡拍的片子用默认的逐句合成;讲解、纪录片这类跟着语义走的旁白,可以整段合成。实测:两段共 3 句,第一段 2 句一次出 11.5 s,第二段单独放到网格上,三句相似度都是 1.00。 +### 逐句拉近语速(`pace`) + +可选的一步。一条 take 里句与句本来就该有快慢,听着顺就不用动;整段合成的 take 偶尔听得出忽快忽慢,这时 `bin/vh pace` 在这条 take 上把各句往同一个语速拉近一部分,不重新合成。什么时候用、为什么只拉一部分,见 `playbook/04-audio.md` 的"一条 take 里快慢差得明显时"。 + +```bash +bin/vh pace projects/

--dry-run # 只量不写:每句的语速、倍数、处理后的语速、句后的停顿 +bin/vh pace projects/

# 往各句的中位数拉近,写回配音和 timeline,再重做字幕 +bin/vh pace projects/

--ref o1-o2 # 参照人听过、认可的几句(一个 id,或"首-尾");--target 6 直接给数 +bin/vh pace projects/

--restore # 放回原来的 take +``` + +- **输入**:`voiceover..wav` 和带 `words` 的 `timeline..json`(`--align gemini` 或 `minimax` 给的);没有逐字时间的,加 `--align gemini|whisper` 先把原 take 转写一遍。`script.txt` 提供段落(空行)和停顿标签。卡拍(`--beats`)的 timeline 和双人对话不处理:重排会让句子离开拍点,两个声音本来就各有各的语速。 +- **每句**:先在 ASR 给的句间边界附近找到两句之间的那段静音,在那里分开(对齐有时把上一句的结尾和下一句的开头放在同一个时刻,还落在上一句最后一个字里);再按波形定起止,从对齐给的句首往回走到声音静下来为止(ASR 把句首放晚了也不会切掉第一个字),句尾同样往后走。句内长于 `--pause`(0.26 s)的停顿压到 0.26 s,切口两侧各淡 6 ms;稿子里这句写了 ``、`` 或 `<#秒#>` 的,停顿是要求过的,保留。语速 = 字数 ÷ 去掉这些多余停顿后的发声时长;数字按念法算(75 是三个字,1953年是五个),英文按音节。倍数 = (参照 ÷ 这句的语速)的 `--strength` 次方(默认 0.8),限在 `--clamp`(0.92–1.18),用 `atempo` 拉,不变调。和参照差不到 5%(`--tolerance`)的句子、不到 4 个字或不到 0.5 s 的短句不动。 +- **句间停顿**:用 take 自己的(上一句人声结束到下一句人声开始),只把太短、太长的收进按句尾定的范围:逗号或没有标点 0.12–0.40 s,句号 0.25–0.70 s,一段的最后一句 0.45–1.20 s。`--extra id=秒` 给某句之后加(负数是减)。第一句之前(`--lead`)和最后一句之后的部分原样留着。拼接在 numpy 里做,不经过 `anullsrc` 和 `concat`(见下文"手动的响度和合成命令")。 +- **写回**:原件留成 `voiceover..raw.wav` 和 `timeline..raw.json`。再跑一次总是从原件出发,参数不会叠加;`bin/vh tts` 重新合成后,新的 take 就是原件(timeline 的 `pace` 里记着写出那个文件的 sha256,用来分辨)。`timeline.json` 和 `voiceover.wav` 是这个语言的拷贝时一起更新。`pace` 里还有参照、参数和每句的语速、倍数、处理后的语速;各句的 `file` 去掉了(`vo//` 里是原 take 的切片)。最后用默认参数跑 `bin/vh captions

`;原来的字幕折得更窄(竖屏的 `zh-max 11`)时会提示用原来的宽度再跑一次。 +- **逐字时间**:默认 `--align map`,把每个字的时间按切口、压掉的停顿和倍数换算到新 take 上,句首、句尾对到波形上的人声。离线、不花钱、每次结果一样。`--align gemini` 把新 take 整条转写一遍(9 分钟以内一次调用),`--align whisper` 用本地 mlx-whisper(只能在 Apple Silicon 上,模型用 `WHISPER_MODEL`,默认 `mlx-community/whisper-large-v3-turbo`);两种都交给 `tts.py` 的 `align_lines`,更新 `asr` 的相似度和标记,并报出和换算的时间差了多少。转写失败(比如当天额度用完)时,新 take 照样写好、保留换算的时间,命令以 1 退出,之后同一条命令再跑一次就行。 + +实测(2026-10-07): +- 合成的测试 take(6 句、42 个已知起点的"音节",快的 7.3、慢的 4.1 个/秒):音节一个不少;换算的逐字起点和真实起点相差中位数 2 ms,最多 40 ms(`atempo` 在一句之内伸缩得不完全均匀),在 30 fps 的一帧以内; +- 一条 2 分钟、27 句的 Gemini 整稿:句内语速 5.18–7.38 字/秒(标准差 0.50)拉到 5.88–6.83(0.21),13 句的速度没动,总长 129.5 s → 121.4 s。这里的量法只算人声的起止之间、数字按念法算,所以和 `playbook/04` 里手工那次的数字不一样。换算出的句首都在波形上人声起点的 0.04 s 以内(多半早 0.03 s,是句首留的余量);同一条用 Gemini 重转写,句首和换算的差中位数 0.03 s、最多 0.13 s;用 whisper 重转写,句首中位数早 0.19 s,最多早 0.27 s。所以默认用换算,重转写主要用来查有没有字丢了、念错了。Gemini 那次把第一句里几个字的时间排到了后面两句,这两句被标出来,音频本身没有问题(whisper 听这两句是 0.94 和 1.00); +- 同一条用 `bin/vh qa` 查 click:原 take 57 处,处理后 53 处,切口没有带进新的。 + ### 音色(`bin/vh voices`) diff --git a/tools/audio/pace.py b/tools/audio/pace.py new file mode 100644 index 0000000..54073b1 --- /dev/null +++ b/tools/audio/pace.py @@ -0,0 +1,530 @@ +"""Bring the pace of the lines in one narration take closer together, without synthesizing again. + +usage (via bin/vh pace): python tools/audio/pace.py [zh|en] [--ref ID[-ID] | --target N] [--strength 0.8] + [--clamp 0.92,1.18] [--tolerance 0.05] [--pause 0.26] [--extra id=s,…] + [--align map|gemini|whisper] [--dry-run] [--restore] + +Optional and gentle: lines in a take are meant to differ in pace, and nothing needs doing when the take sounds fine. +When it audibly runs fast and slow (a one-request take, --join all, often does), this moves each line part of the way +towards one reference pace and keeps the rest of the natural variation (playbook/04-audio.md, "一条 take 里快慢差得明显时"). +The numbers are a reference; the ear decides. --dry-run prints them and writes nothing. + +Input audio/voiceover..wav + audio/timeline..json with word times (bin/vh tts … --align gemini, or minimax); + audio/script.txt for the blocks (blank lines) and for pause tags. Without word times: --align gemini|whisper + transcribes the take first. A timeline on a beat grid (--beats) or a dialogue is refused: re-pacing would move + lines off the grid, and two voices keep their own pace. +Per line: + 1. cut points from the waveform: two lines split at the silence nearest their ASR boundary (an aligner can put one + line's end and the next one's start at the same instant, inside the first one's last syllable); from the aligned + start, step back while the signal is within 32 dB of the take's speech level (an ASR start placed late by a + misheard first syllable would clip it), and forward from the end; + 2. silences inside the line longer than --pause (0.26 s) are shortened to it, 6 ms fades on both sides of each cut. + A line whose script has , or a <#s#> mark keeps its pauses (they were asked for); + 3. pace = spoken units / voiced seconds after step 2: CJK characters (digits read out: 75 = 七十五, 3; 1953年 = 4), or + syllables in English. Reference: the median of the lines (default), --ref the pooled pace of lines a person + approved (--ref hook, --ref s1-s3), or --target N. Each line moves by tempo (ffmpeg atempo, pitch kept) + x = (reference / pace) ^ --strength, within --clamp; lines within --tolerance of it, and lines under 4 units or + 0.5 s, are left as they are; + 4. pauses between lines are the take's own, measured voice to voice, kept inside a range by the line's end: + comma (or no punctuation) 0.12–0.40 s, full stop 0.25–0.70 s, end of a block 0.45–1.20 s; --extra id=s adds + (or, negative, takes away) after a line. The silence before the first line (--lead) and after the last is kept. + Lines are joined in numpy, never with anullsrc + concat (playbook/04-audio.md, "旁白晚一点进来"). +Output voiceover..wav and timeline..json written back (timeline.json / voiceover.wav too when they are + copies of this language), the originals kept as voiceover..raw.wav / timeline..raw.json, then + bin/vh captions . A second run starts again from the .raw take, so settings never compound; + after a new bin/vh tts run the new take is the original (the timeline's "pace" block records a sha256 of the + paced file to tell them apart). --restore puts the original back and removes the .raw files. + Word times: --align map (default) carries each word through the cuts and the tempo change: offline, free, + deterministic. --align gemini transcribes the new take (one call per ≤ 9 min, GEMINI_API_KEY) and --align whisper + with a local mlx-whisper (Apple Silicon; WHISPER_MODEL, default mlx-community/whisper-large-v3-turbo); both go + through tts.py's align_lines like bin/vh tts --align, refresh asr {similarity, flag} and print how far the + mapped times were (a failed transcription keeps the mapped times, exit 1). A whisper "1953" against a script's + 一九五三 is spread over the heard word. + The timeline's "pace" block: reference, settings, and per line its pace before, the tempo, and after. +""" +import argparse, hashlib, json, math, os, re, shutil, subprocess, sys, tempfile +from pathlib import Path + +sys.path.insert(0, str(Path(__file__).resolve().parent)) +import tts # noqa: E402 parse_script, align_lines, transcribe, asr_record: one alignment for the whole harness +try: + import numpy as np +except ImportError: # -h on a plain python3; bin/vh pace runs it with numpy + np = None + +HOP = 0.01 # envelope step (s) +BELOW = 32.0 # "voice" = within this many dB of the take's speech level (95th percentile of the envelope) +FLOOR_DB = -45.0 # … and never below this +HEAD, TAIL = 0.03, 0.06 # margins kept around the voiced extent of a line (s) +EDGE_FADE, CUT_FADE = 0.010, 0.006 +MIN_UNITS, MIN_VOICED = 4, 0.5 # shorter lines keep their pace: too few units to measure one +SPLIT_NEAR = 0.2 # look this far around the ASR boundary between two lines for the silence between them +GAPS = {"comma": (0.12, 0.40), "stop": (0.25, 0.70), "block": (0.45, 1.20)} # voice-to-voice silence after a line (s) +ASR_PIECE = 540.0 # longest piece per gemini transcription (16 kHz inline audio: ≈ 10 min fits 20 MB) +STOPS = tuple("。!?!?….") +CLOSERS = " ”’」』))》\"'" +DIG = "零一二三四五六七八九" +CJK = re.compile(rf"[{tts._CJK}]") +TOKEN = re.compile(rf"[{tts._CJK}]|\d+(?:\.\d+)?%?年?|[A-Za-z]+(?:['’][A-Za-z]+)*") +ASKED_PAUSE = re.compile(r"<\s*(?:short|long)\s+pause\s*>|<#\s*\d", re.I) + +# ---------- spoken units ---------- + +def zh_int(n, head=True): + """The Chinese reading of an integer (1953 → 一千九百五十三); head: 10–19 read 十… at the start.""" + if n >= 10 ** 8: + r = n % 10 ** 8 + return zh_int(n // 10 ** 8, head) + "亿" + (("零" if r < 10 ** 7 else "") + zh_int(r, False) if r else "") + if n >= 10 ** 4: + r = n % 10 ** 4 + return zh_int(n // 10 ** 4, head) + "万" + (("零" if r < 1000 else "") + zh_int(r, False) if r else "") + s, z = "", False + for p, u in ((1000, "千"), (100, "百"), (10, "十"), (1, "")): + d = n // p % 10 + if d: + s += ("零" if z else "") + DIG[d] + u; z = False + elif s: + z = True + s = s or "零" + return s[1:] if head and s.startswith("一十") else s + +def en_syllables(w): + if w.isupper() and len(w) <= 5: # an acronym is read letter by letter + return len(w) + v = re.findall(r"[aeiouy]+", w.lower()) + n = len(v) - (len(v) > 1 and w.lower().endswith("e") and not w.lower().endswith(("le", "ee", "ie", "ye"))) + return max(1, n) + +def units(text, lang): + """Spoken units of a line: CJK characters, digits as read (zh: 七十五 / 一九五三年; en: ~1.3 per digit), English syllables.""" + n = 0 + for t in TOKEN.findall(text): + if CJK.fullmatch(t): + n += 1 + elif t[0].isdigit(): + num, pct, year = t.rstrip("%年"), "%" in t, t.endswith("年") + whole, _, frac = num.partition(".") + if lang == "zh" or CJK.search(text): + k = len(whole) if year and len(whole) == 4 else len(zh_int(int(whole))) + n += k + (1 + len(frac) if frac else 0) + 3 * pct + year + else: + n += max(1, round(1.3 * len(whole.lstrip("0") or "0"))) + (1 + len(frac) if frac else 0) + 2 * pct + else: + n += en_syllables(t) + return n + +# ---------- audio in and out (ffmpeg, float32 mono) ---------- + +def probe(wav): + out = subprocess.run(["ffprobe", "-v", "error", "-select_streams", "a:0", "-show_entries", "stream=sample_rate,codec_name", + "-of", "json", str(wav)], check=True, capture_output=True, text=True).stdout + st = json.loads(out)["streams"][0] + return int(st["sample_rate"]), st["codec_name"] + +def read_wav(wav, sr): + raw = subprocess.run(["ffmpeg", "-v", "error", "-i", str(wav), "-ac", "1", "-ar", str(sr), "-f", "f32le", "pipe:1"], + check=True, capture_output=True).stdout + return np.frombuffer(raw, dtype=" t0 and (c if c is not None else dur) - t0 > ASR_PIECE: + pieces.append((t0, prev)); t0 = prev + prev = c + pieces.append((t0, dur)) + words = [] + for a, b in pieces: + part = tmp / f"asr-{a:.0f}.wav" + subprocess.run(["ffmpeg", "-v", "error", "-y", "-i", str(wav), "-af", f"atrim=start={a:.3f}:end={b:.3f},asetpts=N/SR/TB", str(part)], check=True) + _, ws = tts.transcribe(part, lang, tmp) + words += [dict(w, start=w["start"] + a, end=w["end"] + a) for w in ws] + return words + +def align_take(segs, wav, engine, lang, tmp, dur): + """Transcribe a whole take and place every line on it (start, end, words, asr) like bin/vh tts --align.""" + texts = [s["text"] for s in segs] + cuts = [(a["end"] + b["start"]) / 2 for a, b in zip(segs, segs[1:]) if "end" in a and "start" in b] + res = tts.align_lines(texts, asr_words(wav, engine, lang, texts, cuts, tmp), dur) + return [{"start": round(r["start"], 3), "end": round(r["end"], 3), "words": r["words"], + "asr": tts.asr_record(r["heard"], r, 0.85)} for r in res] + +# ---------- the edit ---------- + +def gap_kind(seg, nxt, block): + if nxt is not None and block.get(nxt["id"]) != block.get(seg["id"]): + return "block" + return "stop" if seg["text"].rstrip(CLOSERS).endswith(STOPS) else "comma" + +def plan(x, sr, segs, lang, block, asked, a): + """Cut points, inner pauses, pace and tempo of every line (nothing written).""" + n = len(x) // int(HOP * sr) + env = x[: n * int(HOP * sr)].reshape(n, -1) + db = 10 * np.log10(np.mean(env ** 2, axis=1) + 1e-18) + voiced = db > max(FLOOR_DB, float(np.percentile(db, 95)) - BELOW) + frame = lambda t: int(t / HOP + 1e-6) + loud = lambda t: 0 <= frame(t) < n and bool(voiced[frame(t)]) + dur = len(x) / sr + # where one line ends and the next begins: the silence nearest the ASR boundary (an aligner can put a line's end and the + # next one's start at the same instant, at the last voiced frame of the first); lines that touch split at the boundary + splits = [] + for p, q in zip(segs, segs[1:]): + m = (p["end"] + q["start"]) / 2 + k, stop, runs = max(frame(min(p["end"], q["start"]) - SPLIT_NEAR), frame(p["start"])), min(frame(max(p["end"], q["start"]) + SPLIT_NEAR), frame(q["end"]), n), [] + while k < stop: + if voiced[k]: + k += 1; continue + j = k + while j < n and not voiced[j]: + j += 1 + k0 = k + while k0 > 0 and not voiced[k0 - 1]: + k0 -= 1 + if j - k0 >= 3: + runs.append((k0 * HOP, j * HOP)) + k = j + splits.append(min(runs, key=lambda r: max(r[0] - m, m - r[1], 0.0)) if runs else (m, m)) + lines, lo = [], 0.0 + for i, s in enumerate(segs): + wlo = splits[i - 1][1] if i else 0.0 # the voice of this line lies between the silences around it + hi = splits[i][0] if i + 1 < len(segs) else dur + hi = max(hi, wlo + HOP) + u = units(s["text"], lang) + a0 = min(max(s["start"], wlo), hi) + if not u: # "……": nothing to measure, the span as it is + v0, v1 = a0, min(max(s["end"], a0), hi) + else: + v0 = a0 + if loud(v0): + while v0 - HOP >= wlo and loud(v0 - HOP): + v0 -= HOP + else: # an early ASR start: move on to the voice + while v0 + HOP < min(s["end"], hi) and not loud(v0): + v0 += HOP + v1 = min(max(s["end"], v0 + HOP), hi) + if loud(v1): + while v1 + HOP <= hi and loud(v1): + v1 += HOP + else: + while v1 - HOP > v0 and not loud(v1 - HOP): + v1 -= HOP + nxt = sum(splits[i]) / 2 if i + 1 < len(segs) else dur # margins reach into the silence, up to its middle + c0, c1 = max(lo, v0 - HEAD * bool(u)), min(max(nxt, v1), v1 + TAIL * bool(u)) + # quiet runs strictly inside the voiced extent: cut the middle of the long ones (kept when the script asks for them) + keep, runs, k, end, seen = np.ones(int(round(c1 * sr)) - int(round(c0 * sr)), bool), [], frame(v0), min(frame(v1), n), False + while u and k < end: + if voiced[k]: + seen = True; k += 1; continue + j = k + while j < end and not voiced[j]: + j += 1 + if seen and j < end and (j - k) * HOP > a.pause > 0: + runs.append((k * HOP, j * HOP)) + k = j + excess = sum(e - b - a.pause for b, e in runs) + if runs and not asked.get(s["id"]): + for b, e in runs: + keep[int(round((b + a.pause / 2 - c0) * sr)): int(round((e - a.pause / 2 - c0) * sr))] = False + voiced_s = max(v1 - v0 - excess, HOP) + lines.append({"seg": s, "u": u, "c0": c0, "c1": c1, "v0": v0, "v1": v1, "keep": keep, "runs": runs, + "rate": u / voiced_s if u else None, "voiced": voiced_s, "kept_pauses": bool(runs and asked.get(s["id"]))}) + lo = c1 + for L in lines: + L["ok"] = L["u"] >= MIN_UNITS and L["voiced"] >= MIN_VOICED + measurable = [L for L in lines if L["ok"]] + if a.target: + ref, how = a.target, f"--target {a.target:g}" + elif a.ref: + ids = ref_ids(a.ref, segs) + pool = [L for L in lines if L["seg"]["id"] in ids and L["u"]] + ref, how = sum(L["u"] for L in pool) / sum(L["voiced"] for L in pool), f"--ref {a.ref} ({len(pool)} line(s))" + elif measurable: + ref, how = float(np.median([L["rate"] for L in measurable])), f"median of {len(measurable)} line(s)" + else: + ref, how = None, "no line long enough to measure" + for L in lines: + f, why = 1.0, "" + if not L["ok"]: + why = "short" if L["u"] else "no words" + elif ref and abs(math.log(ref / L["rate"])) > math.log(1 + a.tolerance): + f = min(a.clamp[1], max(a.clamp[0], (ref / L["rate"]) ** a.strength)) + elif ref: + why = "close enough" + L["f"], L["why"] = f, why + for i, L in enumerate(lines): # voice-to-voice silence after the line, in the take and in the new one + nxt = lines[i + 1] if i + 1 < len(lines) else None + if nxt: + kind = gap_kind(L["seg"], nxt["seg"], block) + g = nxt["v0"] - L["v1"] + lo_k, hi_k = GAPS[kind] + L["gap"], L["kind"] = max(0.04, min(max(g, lo_k), hi_k) + a.extra.get(L["seg"]["id"], 0.0)), kind + L["gap0"] = g + return lines, ref, how + +def ref_ids(spec, segs): + ids = [s["id"] for s in segs] + if spec in ids: + return {spec} + for m in re.finditer("-", spec): + p, q = spec[:m.start()], spec[m.end():] + if p in ids and q in ids and ids.index(p) <= ids.index(q): + return set(ids[ids.index(p): ids.index(q) + 1]) + sys.exit(f"--ref {spec}: not a line id or a range of them (first-last) in this timeline: {' '.join(ids)}") + +def render(x, sr, lines): + """→ (new audio, per line a function: time in the take → time in the new audio); sets each line's real tempo.""" + out, maps = [x[: int(round(lines[0]["c0"] * sr))]], [] + t = len(out[0]) / sr + for i, L in enumerate(lines): + a, b = int(round(L["c0"] * sr)), int(round(L["c1"] * sr)) + chunk, keep = x[a:b].copy(), L["keep"][: b - a] + fz = int(CUT_FADE * sr) + for e in np.flatnonzero(np.diff(keep.astype(np.int8))): # fade out before each cut, in after it (length unchanged) + if keep[e]: + lo = max(0, e + 1 - fz); chunk[lo:e + 1] *= np.linspace(1, 0, e + 1 - lo) + else: + hi = min(len(chunk), e + 1 + fz); chunk[e + 1:hi] *= np.linspace(0, 1, hi - e - 1) + kept = chunk[keep] + y = atempo(kept, sr, L["f"]) + fe = min(int(EDGE_FADE * sr), len(y) // 2) + if fe: + y[:fe] *= np.linspace(0, 1, fe); y[-fe:] *= np.linspace(1, 0, fe) + ck = np.concatenate([[0], np.cumsum(keep)]) + scale = len(y) / max(len(kept), 1) # the real stretch, so a line's last word ends where its audio does + maps.append(lambda r, ck=ck, a=a, t=t, scale=scale: t + ck[min(max(int(round(r * sr)) - a, 0), len(ck) - 1)] * scale / sr) + L["f_real"] = 1 / scale if len(kept) else 1.0 # atempo's length is off by up to ~1% (silence stretches unevenly) + out.append(y); t += len(y) / sr + if i + 1 < len(lines): # the pause: voice-to-voice gap minus the margins already in both chunks + nxt = lines[i + 1] + margins = (L["c1"] - L["v1"]) / L["f_real"] + (nxt["v0"] - nxt["c0"]) / nxt["f"] + pad = np.zeros(int(round(max(0.0, L["gap"] - margins) * sr))) + out.append(pad); t += len(pad) / sr + out.append(x[int(round(lines[-1]["c1"] * sr)):]) # the take's own tail + return np.concatenate(out), maps + +def paced_segments(lines, maps): + segs = [] + for L, m in zip(lines, maps): + s = {k: v for k, v in L["seg"].items() if k != "file"} # vo//NN.wav are cuts of the original take + if s.get("words"): + s["words"] = [dict(w, start=round(m(w["start"]), 3), end=round(max(m(w["end"]), m(w["start"])), 3)) for w in s["words"]] + w0, w1 = s["words"][0], s["words"][-1] # the voice's own edges where the ASR put a word late (or ended early) + w0["start"] = min(w0["start"], round(m(L["v0"]), 3)); w1["end"] = max(w1["end"], round(m(L["v1"]), 3)) + s["start"], s["end"] = s["words"][0]["start"], s["words"][-1]["end"] + else: + s["start"], s["end"] = round(m(L["c0"]), 3), round(m(L["c1"]), 3) + segs.append(s) + return segs + +# ---------- main ---------- + +def main(): + ap = argparse.ArgumentParser(description="Bring the pace of the lines in one narration take closer together (see the module doc).") + ap.add_argument("project"); ap.add_argument("lang", nargs="?", choices=["zh", "en"], help="default: the latest bin/vh tts run's") + g = ap.add_mutually_exclusive_group() + g.add_argument("--ref", help="line id, or first-last: their pooled pace is the reference (default: the median of the lines)") + g.add_argument("--target", type=float, help="reference pace in units per second (CJK characters, or English syllables)") + ap.add_argument("--strength", type=float, default=0.8, help="part of the way: tempo = (reference / pace) ^ strength (0–1)") + ap.add_argument("--clamp", default="0.92,1.18", help="slowest,fastest tempo of a line") + ap.add_argument("--tolerance", type=float, default=0.05, help="lines within this fraction of the reference are left as they are") + ap.add_argument("--pause", type=float, default=0.26, help="longest silence kept inside a line (s); 0 keeps them all") + ap.add_argument("--extra", default="", help="id=s,… seconds added after a line (negative: taken away)") + ap.add_argument("--align", default="map", choices=["map", "gemini", "whisper"], help="word times of the new take") + ap.add_argument("--dry-run", action="store_true", help="measure and print, write nothing") + ap.add_argument("--restore", action="store_true", help="put the original take and timeline back, remove the .raw files") + a = ap.parse_args() + try: + a.clamp = tuple(float(v) for v in a.clamp.split(",")) + assert len(a.clamp) == 2 and 0.5 <= a.clamp[0] <= 1 <= a.clamp[1] <= 2 + except (ValueError, AssertionError): + ap.error("--clamp wants slowest,fastest with 0.5 ≤ slowest ≤ 1 ≤ fastest ≤ 2 (e.g. 0.92,1.18)") + if not 0 <= a.strength <= 1 or not 0 <= a.tolerance < 1 or a.pause < 0 or (a.target is not None and a.target <= 0): + ap.error("--strength 0–1, --tolerance 0–1, --pause ≥ 0, --target > 0") + try: + a.extra = {k.strip(): float(v) for k, v in (p.split("=", 1) for p in a.extra.split(",") if p.strip())} + except ValueError: + ap.error("--extra wants id=seconds,… (e.g. hook=0.3,s4=-0.1)") + if np is None: + sys.exit("pace.py needs numpy: run it as bin/vh pace") + proj = Path(a.project).resolve(); audio = proj / "audio" + if a.lang is None: + latest = audio / "timeline.json" + a.lang = json.loads(latest.read_text(encoding="utf-8")).get("lang") if latest.exists() else None + a.lang = a.lang or next((l for l in ("zh", "en") if (audio / f"timeline.{l}.json").exists()), "zh") + wav, tlp = audio / f"voiceover.{a.lang}.wav", audio / f"timeline.{a.lang}.json" + raw_wav, raw_tl = audio / f"voiceover.{a.lang}.raw.wav", audio / f"timeline.{a.lang}.raw.json" + latest = audio / "timeline.json" + copies = latest.exists() and json.loads(latest.read_text(encoding="utf-8")).get("lang") == a.lang # tts's copies of this run + rel = lambda p: p.relative_to(proj) + + def captions(old): + sys.stdout.flush() + r = subprocess.run([sys.executable, str(Path(__file__).resolve().parent / "captions.py"), str(proj), "--lang", a.lang]) + narrow = [c for c in old if len(c.get("zh") or []) > 1 and sum(0.5 if ord(ch) < 128 else 1 for ch in "".join(c["zh"])) <= 16] + if narrow: # two lines that fit one at the default width: wrapped narrower before + w = max(sum(0.5 if ord(ch) < 128 else 1 for ch in line) for c in old for line in c.get("zh") or []) + print(f" note: the old captions were wrapped narrower (≤ {math.ceil(w)}); bin/vh captions {a.project} {a.lang} {math.ceil(w)} redoes them that way") + return r.returncode + + old_caps = json.loads((audio / "captions.json").read_text(encoding="utf-8")) if (audio / "captions.json").exists() else [] + if a.restore: + if not (raw_wav.exists() and raw_tl.exists()): + sys.exit(f"--restore: no {rel(raw_wav)} / {rel(raw_tl)} to put back") + cur = json.loads(tlp.read_text(encoding="utf-8")) if tlp.exists() else {} + if not (wav.exists() and cur.get("pace") and sha256(wav) == cur["pace"].get("sha256")): + sys.exit(f"--restore: {rel(wav)} is not the take bin/vh pace wrote (a new bin/vh tts run?), so the .raw files are from an" + f" older take; nothing changed (delete them by hand if they are not needed)") + shutil.move(str(raw_wav), str(wav)); shutil.move(str(raw_tl), str(tlp)) + if copies: + shutil.copyfile(tlp, latest); shutil.copyfile(wav, audio / "voiceover.wav") + print(f"restored the original take: {rel(wav)}, {rel(tlp)}") + sys.exit(captions(old_caps)) + if not (wav.exists() and tlp.exists()): + sys.exit(f"no {rel(wav)} / {rel(tlp)}: make the narration with bin/vh tts first") + tl = json.loads(tlp.read_text(encoding="utf-8")) + if tl.get("pace"): # the paced take from an earlier run: start again from the original + if sha256(wav) != tl["pace"].get("sha256"): + sys.exit(f"{rel(wav)} changed after bin/vh pace wrote it, but {rel(tlp)} is still the paced one: run bin/vh tts" + f" again (it writes both), or put the paced file back") + if not (raw_wav.exists() and raw_tl.exists()): + sys.exit(f"{rel(tlp)} is paced, but {rel(raw_wav)} / {rel(raw_tl)} are missing: run bin/vh tts again") + src_wav, tl = raw_wav, json.loads(raw_tl.read_text(encoding="utf-8")) + else: + src_wav = wav # a new take (or the first run): it becomes the original + if raw_wav.exists() and not a.dry_run: + print(f" {rel(wav)} is a new take: it replaces the old {rel(raw_wav)}") + if tl.get("grid"): + sys.exit("the lines sit on a beat grid (bin/vh tts --beats): re-pacing would move them off it; change the pace when synthesizing") + if tl.get("speakers") or any(s.get("speaker") for s in tl["segments"]): + sys.exit("a dialogue: each voice keeps its own pace, so bin/vh pace leaves it alone") + segs = tl["segments"] + if not segs: + sys.exit(f"{rel(tlp)} has no lines") + + script = audio / "script.txt" + block, asked = {}, {} + if script.exists(): + for s in tts.parse_script(script)[0]: + block[s["id"]] = s["_block"] + asked[s["id"]] = bool(ASKED_PAUSE.search((s["zh"] if a.lang == "zh" else s["en"]) or s["zh"] or s["en"])) + for i in a.extra: + if i not in {s["id"] for s in segs}: + sys.exit(f"--extra {i}: no line with that id") + + sr, codec = probe(src_wav) + x = read_wav(src_wav, sr) + dur = len(x) / sr + with tempfile.TemporaryDirectory() as td: + tmp = Path(td) + missing = [s["id"] for s in segs if units(s["text"], a.lang) and not s.get("words")] + if missing: + if a.align == "map": + sys.exit(f"no word times for {', '.join(missing[:6])}{' …' if len(missing) > 6 else ''}: run bin/vh tts … --align gemini" + f" --resume first, or pass --align gemini|whisper here to transcribe the take") + print(f" {len(missing)} line(s) without word times: transcribing the take ({a.align}) …", flush=True) + if a.align == "gemini" and not os.environ.get("GEMINI_API_KEY"): + sys.exit("set GEMINI_API_KEY (Google AI Studio → Get API key) for --align gemini") + for s, r in zip(segs, align_take(segs, src_wav, a.align, a.lang, tmp, dur)): + s.update(r); s.pop("file", None) + tl["align"] = {"model": a.align, "by": "bin/vh pace"} + lines, ref, how = plan(x, sr, segs, a.lang, block, asked, a) + unit = "chars/s" if a.lang == "zh" else "syll/s" + y, maps = render(x, sr, lines) + new = paced_segments(lines, maps) + print(f"{'line':<8} {'units':>5} {'pace':>5} tempo after pause after (take → now)") + for L in lines: + r = L["rate"] + print(f" {L['seg']['id']:<6} {L['u']:5d} " + (f"{r:5.2f} → ×{L['f']:.3f} → {r * L['f_real']:5.2f}" if r else f"{'':5} {'':6} {'':5}") + + (f" {L['kind']:<5} {L['gap0']:.2f} → {L['gap']:.2f} s" if "gap" in L else "") + + (f" ({L['why']})" if L["why"] else "") + (" pauses kept (asked for in the script)" if L["kept_pauses"] else "") + + (f" {len(L['runs'])} pause(s) cut to {a.pause:g} s" if L["runs"] and not L["kept_pauses"] else "")) + m = [L for L in lines if L["ok"]] + if m: + before = np.array([L["rate"] for L in m]); after = np.array([L["rate"] * L["f_real"] for L in m]) + print(f"{len(lines)} line(s) · pace {before.min():.2f}–{before.max():.2f} {unit} (sd {before.std():.2f}) → {after.min():.2f}–{after.max():.2f}" + f" (sd {after.std():.2f}) · reference {ref:.2f} ({how}) · {dur:.2f} s → {len(y) / sr:.2f} s") + if a.dry_run: + print("dry run: nothing written. The numbers are a reference; listen before keeping it.") + return + if src_wav == wav: # the original files as they are + shutil.copyfile(wav, raw_wav); shutil.copyfile(tlp, raw_tl) + if missing: # word times found for the original: keep them with it + raw_tl.write_text(json.dumps(tl, ensure_ascii=False, indent=2), encoding="utf-8") + part = wav.with_name(wav.name + ".part") + write_wav(y, sr, codec, part) + out = {k: v for k, v in tl.items() if k not in ("segments", "duration")} + out.update(duration=round(probe_len(part), 3), segments=new) + out["pace"] = {"from": str(rel(raw_wav)), "reference": round(ref, 3) if ref else None, "how": how, "unit": unit, + "strength": a.strength, "clamp": list(a.clamp), "tolerance": a.tolerance, "pause": a.pause, "extra": a.extra, + "gaps": GAPS, "align": "map", "sha256": sha256(part), # the engine once its transcription is in + "lines": [{"id": L["seg"]["id"], "units": L["u"], "pace": round(L["rate"], 3) if L["rate"] else None, + "tempo": round(L["f"], 4), "after": round(L["rate"] * L["f_real"], 3) if L["rate"] else None, + **({"note": L["why"]} if L["why"] else {})} for L in lines]} + def save(): + body = json.dumps(out, ensure_ascii=False, indent=2) + tlp.with_name(tlp.name + ".part").write_text(body, encoding="utf-8"); os.replace(tlp.with_name(tlp.name + ".part"), tlp) + if part.exists(): # the timeline first: stopped in between, the next run refuses ("changed after bin/vh pace + os.replace(part, wav) # wrote it") instead of taking the paced file for a new take and overwriting the original + if copies: + latest.write_text(body, encoding="utf-8"); shutil.copyfile(wav, audio / "voiceover.wav") + save() + print(f"→ {rel(wav)} ({out['duration']} s) · {rel(tlp)} · the original: {rel(raw_wav)}, {rel(raw_tl)}") + rc = 0 + if a.align != "map": + if a.align == "gemini" and not os.environ.get("GEMINI_API_KEY"): + sys.exit("set GEMINI_API_KEY for --align gemini: the paced take is written with mapped word times") + try: + fresh = align_take(new, wav, a.align, a.lang, tmp, out["duration"]) + except Exception as e: # quota, network: the mapped word times stay + print(f"align {a.align} failed: {e}. The paced take keeps the mapped word times; re-run with --align {a.align} later") + fresh, rc = None, 1 + if fresh: + d = [abs(s["start"] - r["start"]) for s, r in zip(new, fresh) if s.get("words")] or [0.0] + for s, r in zip(new, fresh): + s.update(r) + out["align"] = {"model": a.align, "by": "bin/vh pace", "min_sim": 0.85}; out["pace"]["align"] = a.align + save() + print(f"align {a.align}: line starts moved by median {np.median(d):.3f} s, max {max(d):.3f} s from the mapped times · similarity " + + " ".join(f"{s['id']}={s['asr']['similarity']:.2f}" for s in new)) + for s in (s for s in new if s["asr"].get("flag")): + print(f" FLAG {s['id']}: {s['asr']['flag']} — heard “{s['asr']['text']}”: listen to that line in {rel(wav)}") + sys.exit(captions(old_caps) or rc) + +def probe_len(wav): + return float(subprocess.run(["ffprobe", "-v", "error", "-show_entries", "format=duration", "-of", "csv=p=0", str(wav)], + check=True, capture_output=True, text=True).stdout) + +if __name__ == "__main__": + main() diff --git a/tools/ci.sh b/tools/ci.sh index fc78b38..0ba06ef 100755 --- a/tools/ci.sh +++ b/tools/ci.sh @@ -821,12 +821,103 @@ PY rm -rf "$t" } +# bin/vh pace on a synthetic take: tone-burst "syllables" with known onsets, lines at 4 to 7 a second, a long pause inside a +# line and one the script asks for, gaps too short and too long, an ASR start placed 0.12 s late, and one line's end and the +# next one's start put at the same instant inside the first one's last syllable (an aligner did that). Checked against the +# known syllables, not against the tool's own numbers +pace_checks() { + local t p a rc + if ! command -v uv >/dev/null || ! command -v ffmpeg >/dev/null; then skip "pace" "needs uv and ffmpeg"; return; fi + t=$(mktemp -d "${TMPDIR:-/tmp}/vh-ci.XXXXXX"); p="$t/p" + python3 - "$p/audio" <<'PY' +import json, math, struct, sys, wave +from pathlib import Path +a = Path(sys.argv[1]); a.mkdir(parents=True); SR = 24000 +lines = [("a1", "一二三四五六七八,", 4.0, None, 0.05), ("a2", "九十百千万亿兆京。", 7.0, None, 1.6), # id, text, per second, + ("b1", "甲乙丙丁戊己庚辛壬癸。", 5.5, (5, 0.7), 0.4), ("b2", "好。", 3.0, None, 0.3), # (after unit k, pause), + ("b3", "子丑寅卯辰巳午未。", 5.0, (4, 0.6), 0.5), ("b4", "东南西北中发白。", 5.6, None, 0.0)] # gap after the line +(a / "script.txt").write_text("@a1 一二三四五六七八,\n@a2 九十百千万亿兆京。\n\n@b1 甲乙丙丁戊己庚辛壬癸。\n@b2 好。\n" + "@b3 子丑寅卯辰巳午未。\n@b4 东南西北中发白。\n", encoding="utf-8") +x, segs = [0.0] * int(0.4 * SR), [] +for k, (sid, text, rate, pause, gap) in enumerate(lines): + chars = [c for c in text if c not in ",。"]; burst = 0.75 / rate; words = [] + for j, c in enumerate(chars): + t0, n, f0 = len(x) / SR, int(burst * SR), 140 + 15 * ((j * 7 + k * 3) % 5) + x += [0.25 * math.sin(math.pi * i / n) ** 0.5 * (math.sin(2 * math.pi * f0 * i / SR) + 0.5 * math.sin(4 * math.pi * f0 * i / SR)) for i in range(n)] + words.append({"w": c + (text[-1] if j == len(chars) - 1 else ""), "start": round(t0, 3), "end": round(len(x) / SR, 3)}) + if j + 1 < len(chars): + x += [0.0] * int((1 / rate - burst + (pause[1] if pause and j + 1 == pause[0] else 0)) * SR) + if sid == "a1": + words[0]["start"] += 0.12 # the ASR heard the first syllable late + if sid == "b2": # b1 ends and b2 starts at the same instant, 20 ms before b1's voice ends + segs[-1]["words"][-1]["end"] = words[0]["start"] = segs[-1]["end"] = round(segs[-1]["end"] - 0.02, 3) + segs.append({"id": sid, "zh": text, "en": "", "text": text, "start": words[0]["start"], "end": words[-1]["end"], "words": words}) + x += [0.0] * int(gap * SR) +x += [0.0] * int(0.5 * SR) +with wave.open(str(a / "voiceover.zh.wav"), "wb") as w: + w.setnchannels(1); w.setsampwidth(2); w.setframerate(SR); w.writeframes(struct.pack(f"<{len(x)}h", *(int(round(v * 32767)) for v in x))) +tl = json.dumps({"provider": "say", "lang": "zh", "duration": round(len(x) / SR, 3), "segments": segs}, ensure_ascii=False) +for f in ("timeline.zh.json", "timeline.json"): + (a / f).write_text(tl, encoding="utf-8") +(a / "voiceover.wav").write_bytes((a / "voiceover.zh.wav").read_bytes()) +PY + cp -R "$p/audio" "$t/orig" + vh pace "$p" --dry-run >/dev/null && diff -r "$p/audio" "$t/orig" >/dev/null && ok "pace --dry-run writes nothing" || bad "pace --dry-run changed the project" + if a=$(vh pace "$p" 2>&1) && a=$(uv run -q --no-project --with numpy python - "$p/audio" "$t/orig" 2>&1 <<'PY' +import json, subprocess, sys +import numpy as np +A, O = sys.argv[1], sys.argv[2]; sr = 24000; bad = [] +y = np.frombuffer(subprocess.run(["ffmpeg", "-v", "error", "-i", f"{A}/voiceover.zh.wav", "-f", "f32le", "-ac", "1", "-ar", str(sr), "-"], + check=True, capture_output=True).stdout, " -40 +ons = (np.flatnonzero(np.diff(on.astype(int)) == 1) + 1) * 0.01 +words = [w for s in tl["segments"] for w in s["words"]] +if len(ons) != 42 or len(words) != 42: bad.append(f"{len(ons)} syllables in the paced take, {len(words)} words (42 each)") +else: + err = np.abs(ons - [w["start"] for w in words]) + if err.max() > 0.06: bad.append(f"a mapped word {err.max():.3f} s from its syllable ({words[int(err.argmax())]['w']})") +L = {l["id"]: l for l in P["lines"]} +paced = [l for l in P["lines"] if l.get("note") != "short"] +sd = lambda k: float(np.std([l[k] for l in paced])) +if not sd("after") < 0.8 * sd("pace"): bad.append(f"spread of paces {sd('pace'):.2f} → {sd('after'):.2f}") +if not all(0.92 - 1e-9 <= l["tempo"] <= 1.18 + 1e-9 for l in P["lines"]) or L["b2"]["tempo"] != 1.0 or L["a1"]["tempo"] <= 1 or L["a2"]["tempo"] >= 1: + bad.append(f"tempos {[(l['id'], l['tempo']) for l in P['lines']]}") +def longest_quiet(s): + q = on[int(s["start"] / 0.01): int(s["end"] / 0.01)]; best = cur = 0 + for v in q: + cur = 0 if v else cur + 1; best = max(best, cur) + return best * 0.01 +if longest_quiet(S["b1"]) > 0.3 or longest_quiet(S["b3"]) < 0.5: bad.append(f"inner pauses: b1 {longest_quiet(S['b1']):.2f} (cap 0.26), b3 {longest_quiet(S['b3']):.2f} (asked for)") +gap = lambda i, j: S[j]["start"] - S[i]["end"] +if not (0.1 <= gap("a1", "a2") <= 0.2 and 1.1 <= gap("a2", "b1") <= 1.25 and 0.35 <= gap("b1", "b2") <= 0.45): + bad.append(f"gaps a1-a2 {gap('a1', 'a2'):.2f} (0.12), a2-b1 {gap('a2', 'b1'):.2f} (1.2), b1-b2 {gap('b1', 'b2'):.2f} (the take's 0.4)") +caps = json.load(open(f"{A}/captions.json", encoding="utf-8")) +if [c["start"] for c in caps] != [s["start"] for s in tl["segments"]]: bad.append("captions.json is not from the paced timeline") +print("\n".join(bad)); sys.exit(1 if bad else 0) +PY + ); then ok "pace: every syllable kept and every mapped word on it, paces closer within the clamp, an inner pause capped and an asked-for one kept, gaps in range, originals kept, captions redone" + else bad "pace on a synthetic take: $a"; fi + cp "$p/audio/voiceover.zh.wav" "$t/once.wav" + vh pace "$p" >/dev/null && cmp -s "$t/once.wav" "$p/audio/voiceover.zh.wav" && ok "pace runs again from the original take: the same bytes" || bad "pace: a second run gave a different take" + vh pace "$p" --restore >/dev/null && cmp -s "$p/audio/voiceover.zh.wav" "$t/orig/voiceover.zh.wav" && cmp -s "$p/audio/timeline.zh.json" "$t/orig/timeline.zh.json" \ + && [ ! -e "$p/audio/voiceover.zh.raw.wav" ] && ok "pace --restore puts the original back" || bad "pace --restore" + python3 -c "import json, sys; t = json.load(open(sys.argv[1])); t['grid'] = {'beats': 'm.json'}; json.dump(t, open(sys.argv[1], 'w'))" "$p/audio/timeline.zh.json" + a=$(vh pace "$p" 2>&1); rc=$? + case "$rc:$a" in 1:*"beat grid"*) ok "pace refuses a timeline on a beat grid" ;; *) bad "pace on a beat grid: exit $rc: $a" ;; esac + rm -rf "$t" +} + case "${1:-all}" in - --smoke) smoke_checks; decision_checks; desk_checks; textcheck_checks ;; + --smoke) smoke_checks; pace_checks; decision_checks; desk_checks; textcheck_checks ;; --committed) # exactly what a push would send: HEAD in a clean temporary checkout, uncommitted changes left out w="$(mktemp -d "${TMPDIR:-/tmp}/vh-ci.XXXXXX")/head"; git worktree add -q --detach "$w" HEAD || exit 2 (cd "$w" && tools/ci.sh); rc=$?; git worktree remove --force "$w"; rmdir "$(dirname "$w")"; exit $rc ;; - all) static_checks; doc_checks; smoke_checks; decision_checks; desk_checks; textcheck_checks ;; + all) static_checks; doc_checks; smoke_checks; pace_checks; decision_checks; desk_checks; textcheck_checks ;; *) echo "usage: tools/ci.sh [--smoke | --committed]"; exit 2 ;; esac [ $fails = 0 ] && ok "all checks passed" || printf '\033[31m%s check(s) failed\033[0m\n' "$fails"