Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 9 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,15 @@

## Unreleased

**`bin/vh tts … minimax`: MiniMax narration in one request, timed by its own subtitles**
- Why: a film's narration was made with MiniMax T2A (`speech-2.8-hd`) through a project script. Sending the whole script in one request kept the voice from drifting between blocks, which the maintainer had heard and disliked with block-by-block synthesis, and the request's own subtitles gave every character's time without an ASR pass. This makes that a provider.
- `minimax` posts to `https://api.minimax.io/v1/t2a_v2` (the China host `api.minimaxi.com` rejects international keys with 2049; `MINIMAX_BASE_URL` for a China-platform key). Key from `MINIMAX_TOKEN_PLAN_KEY` (Token Plan, spends its credits) first, then `MINIMAX_API_KEY` (pay-as-you-go); an error names the variable, never the key, with a hint for 1008 (no balance on that billing) and 1004/2049 (wrong host for the key). `MINIMAX_TTS_MODEL` / `_SPEED` / `_VOL` / `_PITCH` are recorded in `_run.json`, so `--resume` refuses a change; speed, volume and pitch are checked against the API's ranges before anything is deleted, and so is the 10,000-character limit of each request. `--instruct` / `[direction]` count only as one of MiniMax's nine emotion words; other text is ignored and that is said once. Default voices: zh `Chinese (Mandarin)_Reliable_Executive`, en `English_expressive_narrator`.
- Word timing from `subtitle_type: word`: the entries are grouped back into the words they were read from ("1953" spans the seven syllables of 一千九百五十三, a `@pronounce` name its pinyin, "Claude" comes as C / la / ude) and found in order in the text sent, then cut into the same script-spelled units elevenlabs gives (the file's character offsets are not used: one response counted the pause marks in them, another did not). Per-line runs get `words`, and `--join block|all` places the lines by them, so a joined minimax run needs no `--align gemini` and no `GEMINI_API_KEY`. `--align gemini` still works: one transcription per block for the similarity check, the times stay MiniMax's. The words are kept beside the audio (`NN.words.json`, `_blockNN.words.json`); `--resume` counts a line or block as done only with them.
- Pause marks `<#0.4#>` (seconds): minimax voices them; captions and every other provider, gemini included, drop them. A mark at a line's edge adds to the pause between it and the next line: inside a joined request to the `<#--gap#>` mark that links them, between requests to the silence. One before the first line or after the last is dropped, and a run of marks is merged, as MiniMax requires. In a joined minimax run `--gap` is part of the text, so `--resume` refuses a change.
- `@pronounce 硖合/(xia2)(he2)` lines in script.txt become `pronunciation_dict.tone` (minimax only; other providers say they read the words as written). A changed dictionary makes `--resume` refuse. A `@pronounce` line without a `/` is narration with the id `pronounce`, as before.
- Docs: tools/audio/README (provider table; a "MiniMax" section on keys and hosts, voices and how to list them with `POST /v1/get_voice`, settings, emotion words, word timing, one request for the whole script, pause marks, `@pronounce`, billing, measurements, and the year pitfall: with `text_normalization` on, "1953 年" is read 一千九百五十三年, so write 一九五三年 in the spoken text; 2001 / 2003 read 两千零一 / 两千零三, which is acceptable); playbook/04 (commands, which provider, who can act, tags, 写念法, the selection table); `bin/vh` help and doctor; type 02; captions.py. Voice ids with a space (MiniMax's Chinese ones) cannot go into `@speakers`, so a minimax dialogue is not possible yet.
- Tested against the API on 2026-10-07: per line, `--join block`, `--join all`, Chinese and English, an emotion word, `--resume` (a deleted block and a missing words file are synthesized again, a changed speed or `@pronounce` is refused), the 1008 and 2049 hints, no key, a speed out of range. The other providers are unchanged: `say` per line (zh and en, with `--beats --snap downbeat`, with `--gap` and `--lead`) and with `--join block` (from the same cached transcription, with and without `--beats`, and with transcription failing) give byte-identical timelines and audio with the old and new tts.py. An independent review then found edge marks dropped between requests, spoken `|x|` missing from the words outside a dialogue, the size limit checked only after the old take was deleted, `--gap` changeable on `--resume`, a space added between a CJK character and a Latin word, and `@pronounce` taking over a line id; all fixed and re-tested against the API.

**Long 4K renders and disk space, canvas density, a voice padded to 8 bits, and qa warns about a stale mix**
- Why: lessons from a 457 s HyperFrames canvas film with narration. Its 4K render refused to start (about 57 GB of temporary frames, 53 GB free); a per-frame `setTransform(1, 0, 0, 1, 0, 0)` dropped the 2× scale; padding the voice with `anullsrc` and the `concat` filter quantised it to 8 bits (2,000+ qa click warnings); `bin/vh mix` stopped on a `ding` given a `pitch`, and `bin/vh qa` then reported on the previous mix.wav without a hint.
- engines/README "出 4K": at 4K every capture is a screenshot (drawElement and Linux BeginFrame capture do not supersample), and multi-worker screenshot capture stores every frame as a JPEG beside the output first, an estimated 4 MB a frame; HyperFrames 0.8.82 refuses to start when its estimate is over 90 % of that disk's free space, and checks again from the first 10 frames' real sizes. Bitrate and render time vary a lot with the picture (this canvas film: about 26 Mbps and 1.4 s per second of film; the 3D intro film: about 125 Mbps and 5–8 s), so the "re-encode 4K over about 135 s" line holds for the intro film only. `HF_CAPTURE_PARALLEL_STREAM=true` streams the frames into the encoder (the 13,710 frames took about 11 minutes with 4 workers on an M3 Max), or `--workers 1`; `NODE_OPTIONS=--max-old-space-size=8192` is HyperFrames' own advice when its worker count may outgrow the heap. In the table: reset a canvas transform with `setTransform(dpr, 0, 0, dpr, 0, 0)`, not `(1, 0, 0, 1, 0, 0)` or `resetTransform()`; a WebGL layer drawn into the canvas sizes its canvas and viewport by the same dpr and computes positions in CSS pixels rather than from `gl_FragCoord`.
Expand Down
4 changes: 2 additions & 2 deletions bin/vh
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@
# (中文 || English; [direction] per line; lines snap to beats; @speakers A=Kore B=Puck + "A: …" lines = a dialogue;
# --align gemini: word timestamps + ASR-vs-script check per line; --join: several lines per request;
# --resume: continue the last run after a failed synthesis or transcription, same script and command)
# providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite
# providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite | minimax
# voices list [lang] | design "<description>" [--out f.wav] | delete <voice_id> Gemini voice library / voice design
# captions <project> [zh|en] [zh-max] [--media-end S] SRT (zh / en / bilingual) + captions.json from the TTS timeline; a caption under 1.8 s is
# lengthened into the gap after it, up to the picture's end (--media-end, else index.html's data-duration, media/final.mp4, the timeline)
Expand Down Expand Up @@ -181,7 +181,7 @@ doctor_downloads() { # what the first runs download on their own and how big it
hub=${HF_HUB_CACHE:-${HF_HOME:-$HOME/.cache/huggingface}/hub}
snap=$(ls -d "$hub/models--${model//\//--}/snapshots/"*/ 2>/dev/null | head -1 || true)
if [ "$(uname)" != Darwin ] || [ "$(uname -m)" != arm64 ]; then
echo "- Qwen3-TTS (bin/vh tts's default voice) runs on Apple Silicon only (mlx-audio): use edge, gemini, dashscope or elevenlabs here"
echo "- Qwen3-TTS (bin/vh tts's default voice) runs on Apple Silicon only (mlx-audio): use edge, gemini, minimax, dashscope or elevenlabs here"
else
command -v uv >/dev/null 2>&1 && ! uv run -q --no-project --offline --with mlx-audio python -c "" >/dev/null 2>&1 && voice_pkgs=1
if [ -n "$snap" ] && [ -e "$(set -- "$snap"*.safetensors; echo "$1")" ]; then echo "$c_ok Qwen3-TTS model $model · $(size_of "$snap")"
Expand Down
11 changes: 7 additions & 4 deletions playbook/04-audio.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@
bin/vh tts projects/<p> qwen Serena zh # 本地开源 Qwen3-TTS(默认;首次下载约 2GB)
bin/vh tts projects/<p> qwen Aiden en # 同一份稿子,英文旁白
bin/vh tts projects/<p> gemini Kore zh --align gemini # 云端 Gemini;--align:逐句词级时间 + ASR 对稿,跑偏的句子会被标出来
bin/vh tts projects/<p> minimax "Chinese (Mandarin)_Reliable_Executive" zh --join all # 云端 MiniMax:整份稿一个请求,音色不漂,自带逐字时间
bin/vh captions projects/<p> zh # → captions.zh/en/bi.srt + captions.json(给引擎画进画面)
bin/vh voices list en-GB --gender female # Gemini 音色库;bin/vh voices design "<描述>" 设计一个音色 → voice_… id
# 配乐:代码作曲,段落对齐镜头,节拍精确
Expand All @@ -39,7 +40,7 @@ bin/vh beats <任意音乐文件> # 外来音乐的节
- 双语写成 `中文 || English`,`--lang` 决定念哪一边,两边都会进时间表;
- 只写一边的句子照原样念。进字幕时,含中日韩字符的算中文,不含的算英文,所以纯英文稿出的是 `captions.en.srt`。想让一句纯英文(比如品牌名)留在中文一侧,写成 `Claude Code ||`;只有标点的一句(`……`)跟着前面的单边句走;
- 可以加 `@id` 前缀,比如 `@hook 一句话,做出一支片子。 || One sentence in, one film out.`。
- **写念法,不写字面**:单位、缩写和符号,TTS 常照字母或字面念。showcase 02 的"2 GHz"被念成了 G、H、Z 三个字母,维护者听出来才重录。送去合成的文字要写真实的念法("2 G赫兹""50 千赫兹"),数字、公式、英文缩写也一样。
- **写念法,不写字面**:单位、缩写和符号,TTS 常照字母或字面念。showcase 02 的"2 GHz"被念成了 G、H、Z 三个字母,维护者听出来才重录。送去合成的文字要写真实的念法("2 G赫兹""50 千赫兹"),数字、公式、英文缩写也一样。年份也是:MiniMax 把"1953 年"念成"一千九百五十三年",要写成"一九五三年"。
- `script.txt` 一句只有一份文本,合成和字幕都用它。念法和字幕必须不同(比如画面烧着"2 GHz")时,就用念法单独合成那一句,timeline 里的字幕文字留原文,再在 `script.txt` 的注释里记下念法(02 就是这样做的)。
- 合成后,单位和数字要人耳听一遍:Gemini 转写的第一遍对这类词也不可靠(02 的新 take 第一遍只有 0.10),不能只看相似度。
- 每句合成后,provider 自带的句首句尾静音会被裁掉,所以句间距离就是 `--gap`,按拍落点时人声也落在拍上(裁法和阈值见 `tools/audio/README.md`;想保留原样,加 `--keep-edges`);
Expand All @@ -59,6 +60,7 @@ bin/vh beats <任意音乐文件> # 外来音乐的节
|---|---|
| `qwen`(默认) | 本地开源 Qwen3-TTS,免费、离线,首次下载约 2 GB;0.6B 模型念英文短句偶尔停不下来,每条都要查 |
| `gemini` | 云端,表演力最好,能用一句话导演语气,适合讲解旁白和双人对话;要 `GEMINI_API_KEY` |
| `minimax` | 云端 MiniMax(`speech-2.8-hd`),中文音色多(`Chinese (Mandarin)_*`)。`--join all` 整份稿一个请求,从头到尾一个声音;自带逐字时间,不用转写;支持停顿标记 `<#0.4#>` 和发音词典 `@pronounce`。语气只认一个情绪词。要 `MINIMAX_TOKEN_PLAN_KEY` 或 `MINIMAX_API_KEY`,细节见 `tools/audio/README.md` 的"MiniMax" |
| `say` / `edge` | 打草稿;只能念,不能演 |
| `elevenlabs` / `dashscope` | 接口已留,还没实测 |

Expand Down Expand Up @@ -184,10 +186,11 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词
- 写情绪、语速、重音、句尾走向、停顿,例如"像在跟朋友分享一个惊人的发现,语速偏快,'一句话'重读,尾音上扬成问句";
- 用"像在……"打比方,比"专业""自然"这种抽象词好用得多;
- 同一段里相邻两句的语气要有落差:问句接答句,铺垫接爆点,快接慢。
- **标签**:`<short pause>`、`<long pause>`、`<breath>`、`<laugh>`、`<sigh>` 可以直接写进句子里。只有 `gemini` 会演出来,其他 provider 和字幕都会自动去掉(规则见 `tools/audio/README.md`)。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。
- **标签**:`<short pause>`、`<long pause>`、`<breath>`、`<laugh>`、`<sigh>` 可以直接写进句子里。只有 `gemini` 会演出来,其他 provider 和字幕都会自动去掉(规则见 `tools/audio/README.md`)。`minimax` 用的是停顿标记 `<#0.4#>`(秒),同样只有它会停。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。
- **给 gemini 的指示要短**:一句话讲清情绪和节奏就够了。年龄、性别、口音属于音色,不要写进指示;长段的人设和导演笔记容易让音色漂移。
- **谁能演**:
- `gemini` 表演力最好,整体和逐句指示都听;
- `minimax` 只认一个情绪词(`calm`、`happy`、`whisper` 等),写别的会被忽略;
- 本地 `qwen` 的 0.6B 模型不接受指示,要换 1.7B 的 instruct 模型;
- `say` 和 `edge` 只能念。
- **写法示例**(逐句导演,后两句用了下文的对拍写法):
Expand Down Expand Up @@ -272,7 +275,7 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词
| 需求 | 方案 |
|---|---|
| 中文配音,本地 | `mlx-audio` 在 Apple Silicon 上跑 Qwen3-TTS(Apache-2.0,支持方言和声音设计)。要克隆声音用 CosyVoice 或 GPT-SoVITS。 |
| 配音,云端,追求稳定 | ElevenLabs 的 `/v1/text-to-speech/{voice_id}/with-timestamps` 直接返回字符级时间(`bin/vh tts` 拼成词级);讲解旁白要表演力、或者要双人对话时,用 Gemini 3.8 Flash TTS(`bin/vh tts … gemini`,能用一句话导演语气,有免费档);国内可选火山豆包、阿里百炼、MiniMax。 |
| 配音,云端,追求稳定 | ElevenLabs 的 `/v1/text-to-speech/{voice_id}/with-timestamps` 直接返回字符级时间(`bin/vh tts` 拼成词级);讲解旁白要表演力、或者要双人对话时,用 Gemini 3.8 Flash TTS(`bin/vh tts … gemini`,能用一句话导演语气,有免费档);MiniMax 已封装(`bin/vh tts … minimax`,自带逐字时间,整份稿一个请求音色不漂);国内还可选火山豆包、阿里百炼。 |
| 免费、先凑合用 | `edge-tts`(微软中文音色,非官方接口,随时可能失效) |
| 词级时间戳 | 已封装:`bin/vh tts … --align gemini`(Gemini 3.5 Transcribe,词级,0.1 s 步长,顺带对稿)。离线的话,中文用 FunASR(字级,带标点);通用用 whisper.cpp(Mac 上有 Metal 加速);已有讲稿或歌词、只需对齐时用 ctc-forced-aligner |
| 节拍 | librosa 或 beat_this(后者 downbeat 更准)。音乐平缓时,检测出来的 BPM 只是一个强加的节拍器,不能拿来硬切。 |
Expand Down Expand Up @@ -301,7 +304,7 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词
```

- `cues` 对应 `SCRIPT.md` 里的 `{cue}` 标记,用来在某个词出现的那一刻触发画面。
- `words` 由 `bin/vh tts … --align gemini`(或 elevenlabs)写入;对稿结果 `asr` 和对话里的 `speaker` 也在每个 segment 里。
- `words` 由 `bin/vh tts … --align gemini`(或 elevenlabs、minimax)写入;对稿结果 `asr` 和对话里的 `speaker` 也在每个 segment 里。
- ClaudeAnimationBase 用 `src/config.js` 的 `bpm` 和 `offset`,`pulse()` 和 `beatN()` 会自动对齐节拍。
- Remotion 的做法是在 `calculateMetadata` 里读取音频时长来设置帧数;HyperFrames 的做法是 `sync-durations` 用实测时长覆盖估计值。

Expand Down
Loading
Loading