Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,7 +108,7 @@ OpenVideoHarness/
├── CLAUDE.md / AGENTS.md 入口(本文件;AGENTS.md 由 bin/vh sync-agents 生成,内容相同)
├── README.md / README.zh-CN.md 给人看的说明;CONTRIBUTING.md 是改本仓库时的规则;install.sh 一键安装
├── LOCAL.md 本机环境(不入库;模板是 LOCAL.example.md)
├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk
├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · pace · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk
├── tools/ bin/vh 背后的脚本;audio/README.md 是声音命令的参数手册,desk/README.md 是审阅台的读取约定;ci.sh 是仓库自检
├── skills/open-video-harness/ 轻量 skill:在任何目录把做视频的请求引到本仓库
├── video-types/ 9 类视频(09 实验中):工作流、审美、禁止项、Prompt 增量块、自查重点、案例
Expand Down
9 changes: 9 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,15 @@

## Unreleased

**`bin/vh pace`: bring the pace of the lines in one narration take closer, without synthesizing again**
- Why: a one-request Gemini take ran audibly fast and slow from line to line, and a project script evened it out on the take itself (the steps went into playbook/04 in #70). The maintainer asked for a command, and for it to stay gentle ("别太严格这个语速"): it is optional, moves each line part of the way, leaves the lines that are already close, and its numbers are a reference for the ear.
- Per line: lines split at the silence nearest the ASR boundary (an aligner had put one line's end and the next one's start at the same instant, inside the first one's last syllable), then cut points refined from the waveform (an ASR start placed late by a misheard first syllable no longer clips it); pauses inside a line longer than `--pause` 0.26 s shortened to it, unless the script asked for them (`<short pause>`, `<long pause>`, `<#s#>`); pace = spoken units per voiced second (CJK characters with digits as they are read, or English syllables); tempo `(reference / pace) ^ --strength` (0.8) within `--clamp` (0.92–1.18), through ffmpeg `atempo`; lines within `--tolerance` (5%) of the reference, or under 4 units or 0.5 s, keep their pace. The reference is the median of the lines, `--ref id` or `--ref first-last` (a passage a person approved), or `--target N`.
- Pauses between lines stay the take's own, voice to voice, kept inside a range set by how the line ends (comma or none 0.12–0.40 s, full stop 0.25–0.70 s, end of a block 0.45–1.20 s); `--extra id=s` adds or takes away after a line. Joined in numpy, not with `anullsrc` + `concat`.
- Writes `voiceover.<lang>.wav` and `timeline.<lang>.json` back (and the `timeline.json` / `voiceover.wav` copies of that language), keeps the originals as `.raw`, and redoes the captions. A second run starts again from the `.raw` take, so settings never compound; a new `bin/vh tts` take replaces it (told apart by a sha256 in the timeline's `pace` block); `--restore` puts the original back; `--dry-run` only measures. A timeline on a beat grid and a dialogue are refused.
- Word times: `--align map` (default) carries every word through the cuts and the tempo change: offline, deterministic. `--align gemini` (one call per ≤ 9 min of audio) or `--align whisper` (a local mlx-whisper, Apple Silicon) transcribe the new take and place the lines with tts.py's `align_lines`; the same flag transcribes a take that has no word times yet. A failed transcription (a spent daily quota) keeps the paced take with the mapped times and exits 1.
- Measured: a synthetic take of 42 syllables with known onsets at 4.1–7.3 per second lost none, and the mapped word starts were 2 ms (median) and 40 ms (max) from the real onsets. A 2-minute, 27-line Gemini take went from 5.18–7.38 chars/s (sd 0.50) to 5.88–6.83 (sd 0.21) with 13 lines' tempo untouched, 129.5 s → 121.4 s; every mapped line start was within 0.04 s of the voice onset, a Gemini re-transcription agreed with them within 0.03 s (median), whisper's line starts came 0.19 s early (median); `bin/vh qa` counted 53 click warnings after against 57 before.
- playbook/04's paragraph names the command, tools/audio/README describes it, CLAUDE.md lists it. tools/ci.sh: a smoke test on a synthetic take (tone-burst syllables at different rates, word times known): nothing written by `--dry-run`, every syllable kept and every mapped word on its syllable, the spread of paces smaller, tempos inside the clamp, an inner pause capped and an asked-for one kept, pauses inside their ranges, two lines with one shared ASR boundary split at the silence, the originals kept, a second run byte-identical, `--restore` byte-identical to the original, a beat-grid timeline refused.

**`bin/vh tts --align gemini`: a spent daily quota fails at once instead of looking hung**
- Why: Gemini's 429 for a spent per-day quota can ask for a short `retryDelay` (27 s, 11.8 s in reports on Google's forum), and tts.py took a 429 for the daily quota only when it asked for more than 90 s. Below that it waited and retried up to 5 times a call, 4 calls at a time, so a 2026-10 film's `--align gemini` looked hung (documented in the entry below).
- tts.py reads the quotas a 429 names: when a `google.rpc.QuotaFailure` violation's `quotaId` or `quotaMetric` has `PerDay` in it (`GenerateRequestsPerDayPerProjectPerModel`, also when it is listed with the per-minute quota), the call fails at once, whatever delay it asks for, with a message that names the quota and says `--resume` continues the run once it resets. The take is kept as before. Per-minute 429s are waited out and retried as before, and a 429 that names no daily quota still fails at once only when it asks for more than 90 s. The same applies to `bin/vh voices`.
Expand Down
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,7 +105,7 @@ OpenVideoHarness/
├── CLAUDE.md / AGENTS.md 入口(本文件;AGENTS.md 由 bin/vh sync-agents 生成,内容相同)
├── README.md / README.zh-CN.md 给人看的说明;CONTRIBUTING.md 是改本仓库时的规则;install.sh 一键安装
├── LOCAL.md 本机环境(不入库;模板是 LOCAL.example.md)
├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk
├── bin/vh 命令行:doctor · setup · types · effort · new · style · recipes · hf-init · install-skill · sync-agents · tts · pace · voices · captions · beats · music · sfx · mix · mux · qa · readcheck · textcheck · storyboard · rhythm · cover-preview · sheet · check · gif · review · desk
├── tools/ bin/vh 背后的脚本;audio/README.md 是声音命令的参数手册,desk/README.md 是审阅台的读取约定;ci.sh 是仓库自检
├── skills/open-video-harness/ 轻量 skill:在任何目录把做视频的请求引到本仓库
├── video-types/ 9 类视频(09 实验中):工作流、审美、禁止项、Prompt 增量块、自查重点、案例
Expand Down
11 changes: 10 additions & 1 deletion bin/vh
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,9 @@
# --align gemini: word timestamps + ASR-vs-script check per line; --join: several lines per request;
# --resume: continue the last run after a failed synthesis or transcription, same script and command)
# providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite | minimax
# pace <project> [zh|en] [--ref <id>[-<id>] | --target N] [--strength 0.8] [--clamp 0.92,1.18] [--tolerance 0.05] [--pause 0.26] [--extra id=s,…] [--align map|gemini|whisper] [--dry-run] [--restore]
# optional: bring the pace of the lines in one take closer (part of the way, natural variation kept), no new synthesis;
# --dry-run only measures. The original stays as voiceover.<lang>.raw.wav; captions are redone
# voices list [lang] | design "<description>" [--out f.wav] | delete <voice_id> Gemini voice library / voice design
# captions <project> [zh|en] [zh-max] [--media-end S] SRT (zh / en / bilingual) + captions.json from the TTS timeline; a caption under 1.8 s is
# lengthened into the gap after it, up to the picture's end (--media-end, else index.html's data-duration, media/final.mp4, the timeline)
Expand Down Expand Up @@ -680,6 +683,11 @@ cmd_tts() {
uv run -q --no-project $deps python "$ROOT/tools/audio/tts.py" "$proj" --provider "$prov" ${lang:+--lang "$lang"} ${voice:+--voice "$voice"} ${instruct:+--instruct "$instruct"} "$@"
}
cmd_voices() { [ $# -gt 0 ] || set -- list; python3 "$ROOT/tools/audio/tts.py" voices "$@"; }
cmd_pace() { # numpy for the waveform; mlx-whisper only with --align whisper (Apple Silicon)
local a prev="" deps="--with numpy"
for a in "$@"; do case "$prev:$a" in --align:whisper|*:--align=whisper) deps="$deps --with mlx-whisper" ;; esac; prev=$a; done
uv run -q --no-project $deps python "$ROOT/tools/audio/pace.py" "$@"
}
cmd_captions() { # --media-end S: where a caption too short for the 1.8 s floor may be lengthened to (default: the picture's length)
local u="usage: bin/vh captions <project> [zh|en] [zh-max] [--media-end S]"
local proj=${1:?$u}; shift
Expand Down Expand Up @@ -767,12 +775,13 @@ cmd_usage() { # one command's lines from the header above, with their continuati
case "${2:-}" in -h|--help) # `bin/vh <command> -h` prints its usage instead of running it with -h as a file, project or preset name
case "$1" in captions|readcheck|review|storyboard|rhythm|cover-preview|voices|mix|recipes|sfx|qa) ;; # these print their own, fuller help
tts) cmd_usage tts; echo; exec python3 "$ROOT/tools/audio/tts.py" -h ;; # + every flag, without building the mlx-audio environment first
pace) cmd_usage pace; echo; exec python3 "$ROOT/tools/audio/pace.py" -h ;;
*) cmd_usage "$1" && exit 0 ;; esac ;;
esac

case "${1:-help}" in
doctor) cmd_doctor ;; setup) cmd_setup ;; types) cmd_types ;;
tts) shift; cmd_tts "$@" ;; voices) shift; cmd_voices "$@" ;; captions) shift; cmd_captions "$@" ;; music) shift; cmd_music "$@" ;; sfx) shift; cmd_sfx "$@" ;; beats) shift; cmd_beats "$@" ;; mix) shift; cmd_mix "$@" ;; qa) shift; cmd_qa "$@" ;; mux) shift; cmd_mux "$@" ;;
tts) shift; cmd_tts "$@" ;; pace) shift; cmd_pace "$@" ;; voices) shift; cmd_voices "$@" ;; captions) shift; cmd_captions "$@" ;; music) shift; cmd_music "$@" ;; sfx) shift; cmd_sfx "$@" ;; beats) shift; cmd_beats "$@" ;; mix) shift; cmd_mix "$@" ;; qa) shift; cmd_qa "$@" ;; mux) shift; cmd_mux "$@" ;;
install-skill) shift; cmd_install_skill "$@" ;;
recipes) shift; python3 "$ROOT/tools/recipes.py" "$@" ;;
review) shift; cmd_review "$@" ;;
Expand Down
10 changes: 5 additions & 5 deletions playbook/04-audio.md
Original file line number Diff line number Diff line change
Expand Up @@ -218,13 +218,13 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词

**语速的依据和用法**:表里是句内语速,不算句间停顿。普通话正常朗读、去掉 150 ms 以上的停顿后约 5.2 字/秒:七种语言的朗读语料里,汉语是 5.18 音节/秒,一个汉字一个音节(Pellegrino, Coupé & Marsico 2011,*Language* 87(3),[doi:10.1353/lan.2011.0057](https://doi.org/10.1353/lan.2011.0057))。电话闲聊按总时长算是每分钟 228–247 字,约 3.8–4.1 字/秒(Yuan, Liberman & Cieri 2006,Interspeech,[doi:10.21437/interspeech.2006-204](https://doi.org/10.21437/interspeech.2006-204))。知识科普的 4.5–5.5 围着正常朗读,原理讲解慢一档。倒推稿子字数时乘 0.85,留出句间停顿【综合】:45 s 的知识短视频约 170–210 字,3 分钟的讲解约 540–690 字。TTS 的实际语速看音色:本节第 4 小节比较 `duck_ratio` 用的那段测试旁白(0.25 s 间距,本地 Qwen3-TTS 0.6B 的 Serena,7 句 107 字)句内约 4.1 字/秒。所以字数只用来起稿,时长以 `timeline.json` 的实测为准(硬规则 2)。

**一条 take 里快慢差得明显时,可以逐句拉近**:整稿一次合成、不写逐句指示,语气能统一,句与句的语速却可能差不少。句子之间本来就该有快慢,人听着不别扭就不用动;听得出忽快忽慢时,不用重合成,在这条 take 上逐句处理(还没做成命令):
**一条 take 里快慢差得明显时,可以逐句拉近**:整稿一次合成、不写逐句指示,语气能统一,句与句的语速却可能差不少。句子之间本来就该有快慢,人听着不别扭就不用动;听得出忽快忽慢时,不用重合成,用 `bin/vh pace` 在这条 take 上逐句处理。先加 `--dry-run` 看每句的语速和要拉的倍数,再写:
1. 起止按波形定:从对齐给的句首往回走,走到声音静下来为止(ASR 给的句首不一定准,第一个字听错时常常偏晚,照着剪会切掉它),句尾同样往后走;
2. 句内过长的停顿压到约 0.25 s,切口两侧各淡几毫秒;
3. 每句用 `atempo`(不变调)往同一个语速拉近一些,不必拉齐;参照取人听过、认可的那一段(比如开头)。只校正一部分,例如倍数取(参照 ÷ 这句的语速)的 0.8 次方,限在 0.92–1.18,自然的起伏留着;
4. 按标点和段落重排句间停顿(拼接别用 `anullsrc` 加 `concat`,见上文"旁白晚一点进来"),写回 `voiceover.<lang>.wav`;整条再对齐一次,更新 `timeline.<lang>.json` 的逐字时间,再重跑 `bin/vh captions`。
2. 句内过长的停顿压到约 0.25 s,切口两侧各淡几毫秒;稿子里写了停顿标签的句子不压;
3. 每句用 `atempo`(不变调)往同一个语速拉近一些,不必拉齐。参照默认取各句的中位数;人听过、认可了某一段(比如开头),用 `--ref` 指定它。只校正一部分:倍数取(参照 ÷ 这句的语速)的 0.8 次方,限在 0.92–1.18,差得不多的句子不动,自然的起伏留着;
4. 句间停顿用 take 自己的,只把太短、太长的按句尾标点和段落收进一个范围,个别地方用 `--extra` 加减;拼接别用 `anullsrc` 加 `concat`(见上文"旁白晚一点进来")。写回 `voiceover.<lang>.wav` 和 `timeline.<lang>.json`,逐字时间跟着换算过去,原件留成 `.raw`,字幕跟着重做。

一支 2026-10 的片子,一条 Gemini 整稿各句原来 4.7–6.7 字/秒(第 1、2 步之后量),这样处理后在 5.6–6.1 之间,听起来还是同一条 take。数字只是参考,以人听着顺为准。
一支 2026-10 的片子,一条 Gemini 整稿各句原来 4.7–6.7 字/秒(第 1、2 步之后量),手工这样处理后在 5.6–6.1 之间,听起来还是同一条 take;`bin/vh pace` 的默认值拉得比这轻。数字只是参考,以人听着顺为准。参数、实测和 `--restore` 见 `tools/audio/README.md` 的"逐句拉近语速"。

**要做 A/B**:`studio` 档位下,同一段旁白至少试两种导演方向,把两版都给人听。实测对比见本节开头:同样三句话,只给一个"纪录片旁白"的整体语气时听起来平;逐句导演、再卡上拍之后,才有起伏和节奏。

Expand Down
Loading
Loading