diff --git a/CHANGELOG.md b/CHANGELOG.md index 95a9af6..a662bfc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,15 @@ ## Unreleased +**`bin/vh tts … minimax`: MiniMax narration in one request, timed by its own subtitles** +- Why: a film's narration was made with MiniMax T2A (`speech-2.8-hd`) through a project script. Sending the whole script in one request kept the voice from drifting between blocks, which the maintainer had heard and disliked with block-by-block synthesis, and the request's own subtitles gave every character's time without an ASR pass. This makes that a provider. +- `minimax` posts to `https://api.minimax.io/v1/t2a_v2` (the China host `api.minimaxi.com` rejects international keys with 2049; `MINIMAX_BASE_URL` for a China-platform key). Key from `MINIMAX_TOKEN_PLAN_KEY` (Token Plan, spends its credits) first, then `MINIMAX_API_KEY` (pay-as-you-go); an error names the variable, never the key, with a hint for 1008 (no balance on that billing) and 1004/2049 (wrong host for the key). `MINIMAX_TTS_MODEL` / `_SPEED` / `_VOL` / `_PITCH` are recorded in `_run.json`, so `--resume` refuses a change; speed, volume and pitch are checked against the API's ranges before anything is deleted, and so is the 10,000-character limit of each request. `--instruct` / `[direction]` count only as one of MiniMax's nine emotion words; other text is ignored and that is said once. Default voices: zh `Chinese (Mandarin)_Reliable_Executive`, en `English_expressive_narrator`. +- Word timing from `subtitle_type: word`: the entries are grouped back into the words they were read from ("1953" spans the seven syllables of 一千九百五十三, a `@pronounce` name its pinyin, "Claude" comes as C / la / ude) and found in order in the text sent, then cut into the same script-spelled units elevenlabs gives (the file's character offsets are not used: one response counted the pause marks in them, another did not). Per-line runs get `words`, and `--join block|all` places the lines by them, so a joined minimax run needs no `--align gemini` and no `GEMINI_API_KEY`. `--align gemini` still works: one transcription per block for the similarity check, the times stay MiniMax's. The words are kept beside the audio (`NN.words.json`, `_blockNN.words.json`); `--resume` counts a line or block as done only with them. +- Pause marks `<#0.4#>` (seconds): minimax voices them; captions and every other provider, gemini included, drop them. A mark at a line's edge adds to the pause between it and the next line: inside a joined request to the `<#--gap#>` mark that links them, between requests to the silence. One before the first line or after the last is dropped, and a run of marks is merged, as MiniMax requires. In a joined minimax run `--gap` is part of the text, so `--resume` refuses a change. +- `@pronounce 硖合/(xia2)(he2)` lines in script.txt become `pronunciation_dict.tone` (minimax only; other providers say they read the words as written). A changed dictionary makes `--resume` refuse. A `@pronounce` line without a `/` is narration with the id `pronounce`, as before. +- Docs: tools/audio/README (provider table; a "MiniMax" section on keys and hosts, voices and how to list them with `POST /v1/get_voice`, settings, emotion words, word timing, one request for the whole script, pause marks, `@pronounce`, billing, measurements, and the year pitfall: with `text_normalization` on, "1953 年" is read 一千九百五十三年, so write 一九五三年 in the spoken text; 2001 / 2003 read 两千零一 / 两千零三, which is acceptable); playbook/04 (commands, which provider, who can act, tags, 写念法, the selection table); `bin/vh` help and doctor; type 02; captions.py. Voice ids with a space (MiniMax's Chinese ones) cannot go into `@speakers`, so a minimax dialogue is not possible yet. +- Tested against the API on 2026-10-07: per line, `--join block`, `--join all`, Chinese and English, an emotion word, `--resume` (a deleted block and a missing words file are synthesized again, a changed speed or `@pronounce` is refused), the 1008 and 2049 hints, no key, a speed out of range. The other providers are unchanged: `say` per line (zh and en, with `--beats --snap downbeat`, with `--gap` and `--lead`) and with `--join block` (from the same cached transcription, with and without `--beats`, and with transcription failing) give byte-identical timelines and audio with the old and new tts.py. An independent review then found edge marks dropped between requests, spoken `|x|` missing from the words outside a dialogue, the size limit checked only after the old take was deleted, `--gap` changeable on `--resume`, a space added between a CJK character and a Latin word, and `@pronounce` taking over a line id; all fixed and re-tested against the API. + **Long 4K renders and disk space, canvas density, a voice padded to 8 bits, and qa warns about a stale mix** - Why: lessons from a 457 s HyperFrames canvas film with narration. Its 4K render refused to start (about 57 GB of temporary frames, 53 GB free); a per-frame `setTransform(1, 0, 0, 1, 0, 0)` dropped the 2× scale; padding the voice with `anullsrc` and the `concat` filter quantised it to 8 bits (2,000+ qa click warnings); `bin/vh mix` stopped on a `ding` given a `pitch`, and `bin/vh qa` then reported on the previous mix.wav without a hint. - engines/README "出 4K": at 4K every capture is a screenshot (drawElement and Linux BeginFrame capture do not supersample), and multi-worker screenshot capture stores every frame as a JPEG beside the output first, an estimated 4 MB a frame; HyperFrames 0.8.82 refuses to start when its estimate is over 90 % of that disk's free space, and checks again from the first 10 frames' real sizes. Bitrate and render time vary a lot with the picture (this canvas film: about 26 Mbps and 1.4 s per second of film; the 3D intro film: about 125 Mbps and 5–8 s), so the "re-encode 4K over about 135 s" line holds for the intro film only. `HF_CAPTURE_PARALLEL_STREAM=true` streams the frames into the encoder (the 13,710 frames took about 11 minutes with 4 workers on an M3 Max), or `--workers 1`; `NODE_OPTIONS=--max-old-space-size=8192` is HyperFrames' own advice when its worker count may outgrow the heap. In the table: reset a canvas transform with `setTransform(dpr, 0, 0, dpr, 0, 0)`, not `(1, 0, 0, 1, 0, 0)` or `resetTransform()`; a WebGL layer drawn into the canvas sizes its canvas and viewport by the same dpr and computes positions in CSS pixels rather than from `gl_FragCoord`. diff --git a/bin/vh b/bin/vh index 00e9332..8aea89f 100755 --- a/bin/vh +++ b/bin/vh @@ -32,7 +32,7 @@ # (中文 || English; [direction] per line; lines snap to beats; @speakers A=Kore B=Puck + "A: …" lines = a dialogue; # --align gemini: word timestamps + ASR-vs-script check per line; --join: several lines per request; # --resume: continue the last run after a failed synthesis or transcription, same script and command) -# providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite +# providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite | minimax # voices list [lang] | design "" [--out f.wav] | delete Gemini voice library / voice design # captions [zh|en] [zh-max] [--media-end S] SRT (zh / en / bilingual) + captions.json from the TTS timeline; a caption under 1.8 s is # lengthened into the gap after it, up to the picture's end (--media-end, else index.html's data-duration, media/final.mp4, the timeline) @@ -181,7 +181,7 @@ doctor_downloads() { # what the first runs download on their own and how big it hub=${HF_HUB_CACHE:-${HF_HOME:-$HOME/.cache/huggingface}/hub} snap=$(ls -d "$hub/models--${model//\//--}/snapshots/"*/ 2>/dev/null | head -1 || true) if [ "$(uname)" != Darwin ] || [ "$(uname -m)" != arm64 ]; then - echo "- Qwen3-TTS (bin/vh tts's default voice) runs on Apple Silicon only (mlx-audio): use edge, gemini, dashscope or elevenlabs here" + echo "- Qwen3-TTS (bin/vh tts's default voice) runs on Apple Silicon only (mlx-audio): use edge, gemini, minimax, dashscope or elevenlabs here" else command -v uv >/dev/null 2>&1 && ! uv run -q --no-project --offline --with mlx-audio python -c "" >/dev/null 2>&1 && voice_pkgs=1 if [ -n "$snap" ] && [ -e "$(set -- "$snap"*.safetensors; echo "$1")" ]; then echo "$c_ok Qwen3-TTS model $model · $(size_of "$snap")" diff --git a/playbook/04-audio.md b/playbook/04-audio.md index 487f12b..6135edc 100644 --- a/playbook/04-audio.md +++ b/playbook/04-audio.md @@ -13,6 +13,7 @@ bin/vh tts projects/

qwen Serena zh # 本地开源 Qwen3-TTS(默认;首次下载约 2GB) bin/vh tts projects/

qwen Aiden en # 同一份稿子,英文旁白 bin/vh tts projects/

gemini Kore zh --align gemini # 云端 Gemini;--align:逐句词级时间 + ASR 对稿,跑偏的句子会被标出来 +bin/vh tts projects/

minimax "Chinese (Mandarin)_Reliable_Executive" zh --join all # 云端 MiniMax:整份稿一个请求,音色不漂,自带逐字时间 bin/vh captions projects/

zh # → captions.zh/en/bi.srt + captions.json(给引擎画进画面) bin/vh voices list en-GB --gender female # Gemini 音色库;bin/vh voices design "<描述>" 设计一个音色 → voice_… id # 配乐:代码作曲,段落对齐镜头,节拍精确 @@ -39,7 +40,7 @@ bin/vh beats <任意音乐文件> # 外来音乐的节 - 双语写成 `中文 || English`,`--lang` 决定念哪一边,两边都会进时间表; - 只写一边的句子照原样念。进字幕时,含中日韩字符的算中文,不含的算英文,所以纯英文稿出的是 `captions.en.srt`。想让一句纯英文(比如品牌名)留在中文一侧,写成 `Claude Code ||`;只有标点的一句(`……`)跟着前面的单边句走; - 可以加 `@id` 前缀,比如 `@hook 一句话,做出一支片子。 || One sentence in, one film out.`。 -- **写念法,不写字面**:单位、缩写和符号,TTS 常照字母或字面念。showcase 02 的"2 GHz"被念成了 G、H、Z 三个字母,维护者听出来才重录。送去合成的文字要写真实的念法("2 G赫兹""50 千赫兹"),数字、公式、英文缩写也一样。 +- **写念法,不写字面**:单位、缩写和符号,TTS 常照字母或字面念。showcase 02 的"2 GHz"被念成了 G、H、Z 三个字母,维护者听出来才重录。送去合成的文字要写真实的念法("2 G赫兹""50 千赫兹"),数字、公式、英文缩写也一样。年份也是:MiniMax 把"1953 年"念成"一千九百五十三年",要写成"一九五三年"。 - `script.txt` 一句只有一份文本,合成和字幕都用它。念法和字幕必须不同(比如画面烧着"2 GHz")时,就用念法单独合成那一句,timeline 里的字幕文字留原文,再在 `script.txt` 的注释里记下念法(02 就是这样做的)。 - 合成后,单位和数字要人耳听一遍:Gemini 转写的第一遍对这类词也不可靠(02 的新 take 第一遍只有 0.10),不能只看相似度。 - 每句合成后,provider 自带的句首句尾静音会被裁掉,所以句间距离就是 `--gap`,按拍落点时人声也落在拍上(裁法和阈值见 `tools/audio/README.md`;想保留原样,加 `--keep-edges`); @@ -59,6 +60,7 @@ bin/vh beats <任意音乐文件> # 外来音乐的节 |---|---| | `qwen`(默认) | 本地开源 Qwen3-TTS,免费、离线,首次下载约 2 GB;0.6B 模型念英文短句偶尔停不下来,每条都要查 | | `gemini` | 云端,表演力最好,能用一句话导演语气,适合讲解旁白和双人对话;要 `GEMINI_API_KEY` | +| `minimax` | 云端 MiniMax(`speech-2.8-hd`),中文音色多(`Chinese (Mandarin)_*`)。`--join all` 整份稿一个请求,从头到尾一个声音;自带逐字时间,不用转写;支持停顿标记 `<#0.4#>` 和发音词典 `@pronounce`。语气只认一个情绪词。要 `MINIMAX_TOKEN_PLAN_KEY` 或 `MINIMAX_API_KEY`,细节见 `tools/audio/README.md` 的"MiniMax" | | `say` / `edge` | 打草稿;只能念,不能演 | | `elevenlabs` / `dashscope` | 接口已留,还没实测 | @@ -184,10 +186,11 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词 - 写情绪、语速、重音、句尾走向、停顿,例如"像在跟朋友分享一个惊人的发现,语速偏快,'一句话'重读,尾音上扬成问句"; - 用"像在……"打比方,比"专业""自然"这种抽象词好用得多; - 同一段里相邻两句的语气要有落差:问句接答句,铺垫接爆点,快接慢。 -- **标签**:``、``、``、``、`` 可以直接写进句子里。只有 `gemini` 会演出来,其他 provider 和字幕都会自动去掉(规则见 `tools/audio/README.md`)。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。 +- **标签**:``、``、``、``、`` 可以直接写进句子里。只有 `gemini` 会演出来,其他 provider 和字幕都会自动去掉(规则见 `tools/audio/README.md`)。`minimax` 用的是停顿标记 `<#0.4#>`(秒),同样只有它会停。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。 - **给 gemini 的指示要短**:一句话讲清情绪和节奏就够了。年龄、性别、口音属于音色,不要写进指示;长段的人设和导演笔记容易让音色漂移。 - **谁能演**: - `gemini` 表演力最好,整体和逐句指示都听; + - `minimax` 只认一个情绪词(`calm`、`happy`、`whisper` 等),写别的会被忽略; - 本地 `qwen` 的 0.6B 模型不接受指示,要换 1.7B 的 instruct 模型; - `say` 和 `edge` 只能念。 - **写法示例**(逐句导演,后两句用了下文的对拍写法): @@ -272,7 +275,7 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词 | 需求 | 方案 | |---|---| | 中文配音,本地 | `mlx-audio` 在 Apple Silicon 上跑 Qwen3-TTS(Apache-2.0,支持方言和声音设计)。要克隆声音用 CosyVoice 或 GPT-SoVITS。 | -| 配音,云端,追求稳定 | ElevenLabs 的 `/v1/text-to-speech/{voice_id}/with-timestamps` 直接返回字符级时间(`bin/vh tts` 拼成词级);讲解旁白要表演力、或者要双人对话时,用 Gemini 3.8 Flash TTS(`bin/vh tts … gemini`,能用一句话导演语气,有免费档);国内可选火山豆包、阿里百炼、MiniMax。 | +| 配音,云端,追求稳定 | ElevenLabs 的 `/v1/text-to-speech/{voice_id}/with-timestamps` 直接返回字符级时间(`bin/vh tts` 拼成词级);讲解旁白要表演力、或者要双人对话时,用 Gemini 3.8 Flash TTS(`bin/vh tts … gemini`,能用一句话导演语气,有免费档);MiniMax 已封装(`bin/vh tts … minimax`,自带逐字时间,整份稿一个请求音色不漂);国内还可选火山豆包、阿里百炼。 | | 免费、先凑合用 | `edge-tts`(微软中文音色,非官方接口,随时可能失效) | | 词级时间戳 | 已封装:`bin/vh tts … --align gemini`(Gemini 3.5 Transcribe,词级,0.1 s 步长,顺带对稿)。离线的话,中文用 FunASR(字级,带标点);通用用 whisper.cpp(Mac 上有 Metal 加速);已有讲稿或歌词、只需对齐时用 ctc-forced-aligner | | 节拍 | librosa 或 beat_this(后者 downbeat 更准)。音乐平缓时,检测出来的 BPM 只是一个强加的节拍器,不能拿来硬切。 | @@ -301,7 +304,7 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词 ``` - `cues` 对应 `SCRIPT.md` 里的 `{cue}` 标记,用来在某个词出现的那一刻触发画面。 -- `words` 由 `bin/vh tts … --align gemini`(或 elevenlabs)写入;对稿结果 `asr` 和对话里的 `speaker` 也在每个 segment 里。 +- `words` 由 `bin/vh tts … --align gemini`(或 elevenlabs、minimax)写入;对稿结果 `asr` 和对话里的 `speaker` 也在每个 segment 里。 - ClaudeAnimationBase 用 `src/config.js` 的 `bpm` 和 `offset`,`pulse()` 和 `beatN()` 会自动对齐节拍。 - Remotion 的做法是在 `calculateMetadata` 里读取音频时长来设置帧数;HyperFrames 的做法是 `sync-durations` 用实测时长覆盖估计值。 diff --git a/tools/audio/README.md b/tools/audio/README.md index 6ebdcea..65025c0 100644 --- a/tools/audio/README.md +++ b/tools/audio/README.md @@ -18,6 +18,7 @@ | `dashscope` | 阿里云百炼 Qwen3-TTS(`qwen3-tts-flash`),云端 | `DASHSCOPE_API_KEY` | 句级 | 接口已留,待实测 | | `elevenlabs` | 高质量多语种配音 | `ELEVENLABS_API_KEY`、voice_id | **词级**(接口给字符级,`bin/vh tts` 拼成词) | 接口已留,待实测 | | `gemini` / `gemini-lite` | Google Gemini 3.8 Flash TTS(2026-09-23 发布)/ Flash-Lite TTS,云端。表演力强,适合讲解旁白和双人对话。Flash 支持 130 种语言,Lite 支持 101 种,都含中文,语言按文本自动识别。`--voice` 用 30 个 studio 音色之一(默认 Kore,每个都能说所有支持的语言)、音色库 id 或自己设计的 `voice_…`(见下文"音色");`bin/vh tts` 的第 5 个参数(`--instruct`)用一句话描述语气,例如"平静、笃定的纪录片旁白";稿子里可以直接写 ``、``、`` 这类标签,只有 gemini 会演出来,其他 provider 和字幕都会自动去掉 | `GEMINI_API_KEY`(Google AI Studio)。有免费档,但免费档的内容可能被 Google 用来改进产品;付费价格见下文"费用" | 句级;加 `--align gemini` 后为词级 | ✅ `gemini`:2026-09-30 实测中英(Kore / Charon),一次通过,逗号处有自然停顿;双人对话、整段合成、设计音色同日实测。`gemini-lite` 走同一条代码路径(只换模型 id),没有单独实测 | +| `minimax` | MiniMax T2A v2(`speech-2.8-hd`),国际站 `api.minimax.io`,云端。中文系统音色是 `Chinese (Mandarin)_*`(默认 `Chinese (Mandarin)_Reliable_Executive`),英文默认 `English_expressive_narrator`。句子里可以写停顿标记 `<#0.4#>`,稿子里可以写发音词典 `@pronounce`;`--join all` 整份稿一个请求,音色不漂。见下文"MiniMax" | `MINIMAX_TOKEN_PLAN_KEY`(Token Plan,优先)或 `MINIMAX_API_KEY`(按量付费) | **词级**(请求自带逐字时间,不用转写) | ✅ 2026-10-07 实测中英:逐句、`--join block`、`--join all`、`--resume` | **已知问题**:默认的本地 Qwen3-TTS 0.6B 模型念英文短句时偶尔停不下来。实测 "One sentence in, a film out." 念完后又多出约 10 秒低电平的含糊声音,整句 12.6 s。所以每条配音都要查一遍:加 `--align gemini` 让机器查(见下一节),或者至少看一眼 `timeline..json` 里的时长。被标出来的那句重跑,或者换 1.7B 模型、换 `gemini`。同一稿 Gemini 的四条都正常。 @@ -35,7 +36,7 @@ - 模拟 Qwen 停不下来:同一句后面接 8 s 低音量的无关絮语,相似度掉到 0.47,并报"最后一个词之后还有 5.5 s";接 8 s 近乎无声的底噪,文字全对(1.00),但报"最后一个词之后还有 8.2 s"。两种都抓到了; - 双人对话里,ASR 会把 B 说话时 A 的一声"嗯"也转出来。这种插话按稿子里的 `|嗯|` 算,不扣分。 -局限:时间戳的步长是 0.1 s(30 fps 下 3 帧);数字可能被规范化("二十六"转成"26"),相似度会降,但不算念错;每次调用有 3–7 s 延迟,逐句模式下 4 句并发。费用约 $0.005 / 分钟音频(付费档),免费档不收费。Tier 1 每分钟只能转写 10 次,超过 10 句(算上复查)时会碰到 429:这时按接口给的时间等(通常不到 1 分钟)再重试,12 句实测 155 s 跑完。接口要求等 90 s 以上的是日配额用完了,会直接报错。转写失败到底(日配额、断网、某个文件坏了)也不丢配音:`voiceover..wav` 和 timeline 在合成后就先写好,出错的句子保留实测时长、`asr` 里记 `error`,命令以非 0 退出;修好原因后用**同一条命令**加 `--resume` 接着跑:只补合成缺的文件(合成中途失败也一样),转写成功过的句子不再重转(结果存在 `vo//*.asr.json`),其余重做;`--join` 时整段音频留在 `vo//_blockNN.wav`,在那之前它的几句共用整段的起止。那次运行的稿子、provider、音色、语气指示(`--instruct`)和 `--join` 方式记录在 `vo//_run.json` 里,`--resume` 要求它们一样,改了就拒绝并指出改了什么(对拍、间距这些时间参数可以改)。不带 `--align` 的普通合成中途失败,也用同一条命令加 `--resume` 接着合成;合成好的普通配音之后再加 `--align gemini --resume`,只做转写,不再合成。key 在合成前就检查。 +局限:时间戳的步长是 0.1 s(30 fps 下 3 帧);数字可能被规范化("二十六"转成"26"),相似度会降,但不算念错;每次调用有 3–7 s 延迟,逐句模式下 4 句并发。费用约 $0.005 / 分钟音频(付费档),免费档不收费。Tier 1 每分钟只能转写 10 次,超过 10 句(算上复查)时会碰到 429:这时按接口给的时间等(通常不到 1 分钟)再重试,12 句实测 155 s 跑完。接口要求等 90 s 以上的是日配额用完了,会直接报错。转写失败到底(日配额、断网、某个文件坏了)也不丢配音:`voiceover..wav` 和 timeline 在合成后就先写好,出错的句子保留实测时长、`asr` 里记 `error`,命令以非 0 退出;修好原因后用**同一条命令**加 `--resume` 接着跑:只补合成缺的文件(合成中途失败也一样),转写成功过的句子不再重转(结果存在 `vo//*.asr.json`),其余重做;`--join` 时整段音频留在 `vo//_blockNN.wav`,在那之前它的几句共用整段的起止(`minimax` 例外:句子照样按它自己的逐字时间切开)。那次运行的稿子、provider、音色、语气指示(`--instruct`)和 `--join` 方式记录在 `vo//_run.json` 里,`--resume` 要求它们一样,改了就拒绝并指出改了什么(对拍、间距这些时间参数可以改)。不带 `--align` 的普通合成中途失败,也用同一条命令加 `--resume` 接着合成;合成好的普通配音之后再加 `--align gemini --resume`,只做转写,不再合成。key 在合成前就检查。 需要更细的强制对齐(音素级、离线)时,仍然可以用 mlx-audio 的 Qwen3-ForcedAligner、FunASR 或 whisper.cpp,这些没有封装。 @@ -70,7 +71,7 @@ ### 整段合成(`--join block|all`) -`--join block` 把空行之间的连续几句放进一个请求,`--join all` 把整份稿放进一个请求(gemini 单个请求最多 8,192 个输入 token)。句与句之间的语气是连着的,比一句一句拼起来自然。行边界从词级时间反推,所以会自动打开 `--align gemini`,`vo//NN.wav` 是从整段音频里切出来的。其他 provider 也能用,只是把几句拼成一段文字去念。 +`--join block` 把空行之间的连续几句放进一个请求,`--join all` 把整份稿放进一个请求(gemini 单个请求最多 8,192 个输入 token)。句与句之间的语气是连着的,比一句一句拼起来自然。行边界从词级时间反推,所以会自动打开 `--align gemini`,`vo//NN.wav` 是从整段音频里切出来的。其他 provider 也能用,只是把几句拼成一段文字去念。`minimax` 例外:句间插 `--gap` 秒的停顿标记,行边界用它自己返回的逐字时间,不用转写(见下文"MiniMax")。 代价是没法逐句控制: - 某一句念得不好,只能整段重跑; @@ -98,8 +99,37 @@ - **语气写短**:`--instruct` 和 `[指示]` 最后都进 `speech_metadata.style`。官方建议只写这一句的情境语气("压低声音,卖个关子");年龄、性别、口音这类身份特征不要写进 style,要换音色,或者用 `voices design` 做一个。长篇的"角色设定""导演笔记"是音色漂移最常见的原因; - **每段 Gemini 音频都带 SynthID 水印**(听不出来,但能检测到)。用了 Gemini 旁白的片子,在项目 `NOTES.md` 里写明"旁白为 AI 合成(Gemini TTS,含 SynthID 水印)",发布时按平台要求标注。 +### MiniMax(`minimax`) + +```bash +bin/vh tts projects/

minimax "Chinese (Mandarin)_Reliable_Executive" zh --join all # 整份稿一个请求 +MINIMAX_TTS_SPEED=1.25 bin/vh tts projects/

minimax - zh --join all --gap 0.3 # 默认音色,语速 1.25,句间停 0.3 s +``` + +- **key 和接口**:请求发到国际站 `https://api.minimax.io/v1/t2a_v2`。key 只从环境变量读,先找 `MINIMAX_TOKEN_PLAN_KEY`(Token Plan 的 Subscription Key,花的是买的 Credits),没有再用 `MINIMAX_API_KEY`(按量付费,账户没有余额时报 `1008 insufficient balance`)。买了 Credits 却只设了按量付费的 key,也是 1008。国际站的 key 发到国内站 `api.minimaxi.com` 会报 `2049 invalid api key`,所以默认走国际站;用国内站(platform.minimaxi.com)的 key 时,设 `MINIMAX_BASE_URL=https://api.minimaxi.com`。报错只写 key 来自哪个变量,不写 key 本身; +- **音色**:`--voice` 写系统音色的 id。中文音色都以 `Chinese (Mandarin)_` 开头,例如 `Chinese (Mandarin)_Reliable_Executive`(沉稳的中年男声,默认)、`_News_Anchor`(女声新闻主播)、`_Male_Announcer`、`_Radio_Host`;英文默认 `English_expressive_narrator`。全部系统音色(2026-10-07 有 332 个)用 `POST /v1/get_voice`、body `{"voice_type": "system"}` 列出,`bin/vh voices` 只管 Gemini: + + ```bash + curl -s https://api.minimax.io/v1/get_voice -H "Authorization: Bearer $MINIMAX_TOKEN_PLAN_KEY" -H "Content-Type: application/json" -d '{"voice_type":"system"}' | python3 -c 'import json,sys; [print(v["voice_id"], "|", v.get("voice_name", "")) for v in json.load(sys.stdin)["system_voice"]]' + ``` + + id 里有空格,命令行里要加引号;也因为有空格,写不进 `@speakers`,双人对话暂时用别家; +- **参数**:`MINIMAX_TTS_MODEL`(默认 `speech-2.8-hd`)、`MINIMAX_TTS_SPEED`(语速 0.5–2,默认 1)、`MINIMAX_TTS_VOL`(音量 0.01–10,默认 1)、`MINIMAX_TTS_PITCH`(音调 −12–12,默认 0)。超出范围在合成前就报错。这几个值记在 `vo//_run.json` 里,`--resume` 时和那次不一样就拒绝。`language_boost` 按 `--lang` 给 `Chinese` 或 `English`,`text_normalization` 打开(数字、符号按念法读,见下面的"年份"); +- **语气**:MiniMax 不听一句话的导演指示,只认一个情绪词:`happy`、`sad`、`angry`、`fearful`、`disgusted`、`surprised`、`calm`、`fluent`、`whisper`。`--instruct` 和 `[ ]` 里写这些词才有用(两处都写了情绪词时取 `[ ]` 的),别的文字会被忽略,命令提示一次。整段合成时只用 `--instruct`,逐句的 `[ ]` 不进请求; +- **词级时间不用转写**:请求时打开字幕(`subtitle_type: word`),MiniMax 返回每个字的起止(毫秒)。`bin/vh tts` 把它们拼回稿子的写法:中文一字一个 word,英文一词一个,"1953" 这样的数字是一个 word,覆盖它念出来的全部音节,`@pronounce` 里的词也一样。所以 timeline 里直接有 `words`,`bin/vh captions` 可以逐字出字幕;`--join` 也不会因此打开转写,不要 `GEMINI_API_KEY`。想要 ASR 对稿时照样可以加 `--align gemini`:整段合成时每段转写一次,相似度按句算,句子的起止仍用 MiniMax 的时间; +- **整份稿一个请求,音色不漂**:逐句或逐段分开合成时,段与段之间能听出像换了个人。`--join all` 把整份稿放进一个请求(不超过 10,000 字符,中文旁白约半小时),从头到尾是同一口气,讲解、纪录片这类旁白推荐这样做。句与句之间插一个 `--gap` 秒的停顿标记(默认 0.25 s),再按 MiniMax 的逐字时间把每句切回 `vo//NN.wav`。代价和"整段合成"一节一样:不能单独重念一句,也不能逐句卡拍。这时 `--gap` 已经写进了请求,`--resume` 不能改它(别的时间参数照样可以改);整段音频和它的逐字时间(`_blockNN.words.json`)都在,才算这一段做完。一个请求超过 10,000 字符时,命令在删掉上一次的配音之前就报错; +- **停顿标记**:句子里写 `<#0.4#>`(秒,0.01–99.99),MiniMax 在那里停 0.4 s;字幕和其他 provider 都会去掉它。写在一句开头或末尾的标记,加到这一句和前后一句之间的停顿上:同一个请求里加进连接两句的停顿标记,分开的请求之间加进静音,逐句、分段、整段合成都一样。比如一段讲完想多停一会儿,就在这段最后一句末尾写 `<#0.5#>`。第一句前面和最后一句后面的标记会被丢掉;连着写的几个标记合成一个(MiniMax 不接受连续的标记); +- **发音词典**:人名、多音字念错时,在 `script.txt` 里加一行 `@pronounce 硖合/(xia2)(he2)`(拼音加声调数字,一个字一组括号;英文词可以写 `omg/oh my god`),可以写多行,不带空格的几条也可以写在同一行。没有 `/` 的 `@pronounce …` 仍是一句 id 为 pronounce 的旁白。它们进请求的 `pronunciation_dict.tone`,只对 `minimax` 生效(gemini 没有发音词典,要改写成同音字)。改了词典,`--resume` 会拒绝,要重新合成; +- **年份会被念成数**:开着 `text_normalization`,"1953 年"念成"一千九百五十三年"。送去合成的文字里,年份写成"一九五三年"。2001、2003 念成"两千零一""两千零三",可以接受。`script.txt` 的文字同时是字幕,字幕要显示"1953 年"时,按 `playbook/04-audio.md`"写念法"的办法单独处理那一句; +- **计费**:按字符算(响应里的 `usage_characters`)。实测一个汉字算 2 个,字母、数字、标点、空格和停顿标记各算 1 个; +- **实测**(2026-10-07,`speech-2.8-hd`): + - 三句中文 `--join all`:一个请求,整条命令约 9 s,出 10.3 s 音频,三句的起止和逐字时间都来自字幕;句中 `<#0.4#>` 处实际停了约 0.8 s(加上逗号自己的停顿); + - 逐句合成时,每句前后各有约 0.3 s 空白,会被自动裁掉; + - 一支片子里 `Chinese (Mandarin)_Reliable_Executive` 在语速 1.25 下约 5.35 字/秒(说话时)。 + ### 旁白里的标签 +- **停顿标记**:`<#0.4#>` 是 MiniMax 的停顿(秒),只有 `minimax` 会停,其他 provider 和字幕都会去掉它,见上一节。 - **标签**:``、``、``、``、`` 可以直接写进句子里。只有 `gemini` 会演出来;其他 provider 和字幕都会自动去掉。去掉的规则:这 5 个和 `` 一律去掉;别的 `<词>` 只要不是两边都紧贴字母或数字,也当标签去掉,所以 `xz` 这类式子会原样保留。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。 ## 外来音乐的节拍(`beats`) diff --git a/tools/audio/captions.py b/tools/audio/captions.py index cdc55a4..e976d3c 100644 --- a/tools/audio/captions.py +++ b/tools/audio/captions.py @@ -8,7 +8,7 @@ (HyperFrames / p5 / Manim read it; every caption stays a pure function of t). Soft subtitle tracks for players/platforms: bin/vh mux