From b301e9a9309da9e081b19b944eff1a7a9cd1ce08 Mon Sep 17 00:00:00 2001 From: zhanglinghao Date: Wed, 7 Oct 2026 19:49:31 +0800 Subject: [PATCH 1/2] bin/vh tts: a minimax provider (MiniMax T2A v2), timed by its own subtitles MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A film's narration was made with MiniMax speech-2.8-hd through a project script: the whole script in one request kept the voice from drifting between blocks, and the request's own subtitles gave every character's time without an ASR pass. This makes it a provider. - POST https://api.minimax.io/v1/t2a_v2; key from MINIMAX_TOKEN_PLAN_KEY, then MINIMAX_API_KEY (never printed); hints for 1008 and 1004/2049; MINIMAX_BASE_URL for a China-platform key; MINIMAX_TTS_MODEL / _SPEED / _VOL / _PITCH checked up front and recorded for --resume; --instruct and [direction] count only as one of MiniMax's emotion words. - subtitle_type word → script-spelled words; --join block|all places the lines by them, so a joined minimax run needs no Gemini. The words are kept beside the audio for --resume. - Pause marks <#0.4#> are voiced by minimax and dropped everywhere else; @pronounce lines become pronunciation_dict.tone. - Docs: tools/audio/README (provider table, a MiniMax section with the year pitfall), playbook/04, bin/vh help and doctor, type 02, CHANGELOG. - Other providers unchanged: say per line (zh, en) and with --join block give byte-identical timelines and audio with the old and new tts.py. --- CHANGELOG.md | 9 + bin/vh | 4 +- playbook/04-audio.md | 11 +- tools/audio/README.md | 32 ++- tools/audio/captions.py | 2 +- tools/audio/tts.py | 313 +++++++++++++++++++++++++++--- video-types/02-knowledge-short.md | 2 +- 7 files changed, 335 insertions(+), 38 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 95a9af6..2860b26 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,15 @@ ## Unreleased +**`bin/vh tts … minimax`: MiniMax narration in one request, timed by its own subtitles** +- Why: a film's narration was made with MiniMax T2A (`speech-2.8-hd`) through a project script. Sending the whole script in one request kept the voice from drifting between blocks, which the maintainer had heard and disliked with block-by-block synthesis, and the request's own subtitles gave every character's time without an ASR pass. This makes that a provider. +- `minimax` posts to `https://api.minimax.io/v1/t2a_v2` (the China host `api.minimaxi.com` rejects international keys with 2049; `MINIMAX_BASE_URL` for a China-platform key). Key from `MINIMAX_TOKEN_PLAN_KEY` (Token Plan, spends its credits) first, then `MINIMAX_API_KEY` (pay-as-you-go); an error names the variable, never the key, with a hint for 1008 (no balance on that billing) and 1004/2049 (wrong host for the key). `MINIMAX_TTS_MODEL` / `_SPEED` / `_VOL` / `_PITCH` are checked against the API's ranges before anything is deleted and recorded in `_run.json`, so `--resume` refuses a change. `--instruct` / `[direction]` count only as one of MiniMax's nine emotion words; other text is ignored and that is said once. Default voices: zh `Chinese (Mandarin)_Reliable_Executive`, en `English_expressive_narrator`. +- Word timing from `subtitle_type: word`: the entries are grouped back into the words they were read from ("1953" spans the seven syllables of 一千九百五十三, a `@pronounce` name its pinyin, "Claude" comes as C / la / ude) and found in order in the text sent, then cut into the same script-spelled units elevenlabs gives (the file's character offsets are not used: one response counted the pause marks in them, another did not). Per-line runs get `words`, and `--join block|all` places the lines by them, so a joined minimax run needs no `--align gemini` and no `GEMINI_API_KEY`. `--align gemini` still works: one transcription per block for the similarity check, the times stay MiniMax's. The words are kept beside the audio (`NN.words.json`, `_blockNN.words.json`); `--resume` counts a line or block as done only with them. +- Pause marks `<#0.4#>` (seconds): minimax voices them; captions and every other provider, gemini included, drop them. In a joined request the lines are linked by a `<#--gap#>` mark plus any mark written at a line's edge; a mark at the edge of a request is dropped and a run of marks merged, as MiniMax requires. +- `@pronounce 硖合/(xia2)(he2)` lines in script.txt become `pronunciation_dict.tone` (minimax only; other providers say they read the words as written). A changed dictionary makes `--resume` refuse. +- Docs: tools/audio/README (provider table; a "MiniMax" section on keys and hosts, voices and how to list them with `POST /v1/get_voice`, settings, emotion words, word timing, one request for the whole script, pause marks, `@pronounce`, billing, measurements, and the year pitfall: with `text_normalization` on, "1953 年" is read 一千九百五十三年, so write 一九五三年 in the spoken text; 2001 / 2003 read 两千零一 / 两千零三, which is acceptable); playbook/04 (commands, which provider, who can act, tags, 写念法, the selection table); `bin/vh` help and doctor; type 02; captions.py. Voice ids with a space (MiniMax's Chinese ones) cannot go into `@speakers`, so a minimax dialogue is not possible yet. +- Tested against the API on 2026-10-07: per line, `--join block`, `--join all`, Chinese and English, an emotion word, `--resume` (a deleted block and a missing words file are synthesized again, a changed speed or `@pronounce` is refused), the 1008 and 2049 hints, no key, a speed out of range. The other providers are unchanged: `say` per line (zh and en) and with `--join block` (from the same cached transcription, and with transcription failing) give byte-identical timelines and audio with the old and new tts.py. + **Long 4K renders and disk space, canvas density, a voice padded to 8 bits, and qa warns about a stale mix** - Why: lessons from a 457 s HyperFrames canvas film with narration. Its 4K render refused to start (about 57 GB of temporary frames, 53 GB free); a per-frame `setTransform(1, 0, 0, 1, 0, 0)` dropped the 2× scale; padding the voice with `anullsrc` and the `concat` filter quantised it to 8 bits (2,000+ qa click warnings); `bin/vh mix` stopped on a `ding` given a `pitch`, and `bin/vh qa` then reported on the previous mix.wav without a hint. - engines/README "出 4K": at 4K every capture is a screenshot (drawElement and Linux BeginFrame capture do not supersample), and multi-worker screenshot capture stores every frame as a JPEG beside the output first, an estimated 4 MB a frame; HyperFrames 0.8.82 refuses to start when its estimate is over 90 % of that disk's free space, and checks again from the first 10 frames' real sizes. Bitrate and render time vary a lot with the picture (this canvas film: about 26 Mbps and 1.4 s per second of film; the 3D intro film: about 125 Mbps and 5–8 s), so the "re-encode 4K over about 135 s" line holds for the intro film only. `HF_CAPTURE_PARALLEL_STREAM=true` streams the frames into the encoder (the 13,710 frames took about 11 minutes with 4 workers on an M3 Max), or `--workers 1`; `NODE_OPTIONS=--max-old-space-size=8192` is HyperFrames' own advice when its worker count may outgrow the heap. In the table: reset a canvas transform with `setTransform(dpr, 0, 0, dpr, 0, 0)`, not `(1, 0, 0, 1, 0, 0)` or `resetTransform()`; a WebGL layer drawn into the canvas sizes its canvas and viewport by the same dpr and computes positions in CSS pixels rather than from `gl_FragCoord`. diff --git a/bin/vh b/bin/vh index 00e9332..8aea89f 100755 --- a/bin/vh +++ b/bin/vh @@ -32,7 +32,7 @@ # (中文 || English; [direction] per line; lines snap to beats; @speakers A=Kore B=Puck + "A: …" lines = a dialogue; # --align gemini: word timestamps + ASR-vs-script check per line; --join: several lines per request; # --resume: continue the last run after a failed synthesis or transcription, same script and command) -# providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite +# providers: qwen (default, local open-source Qwen3-TTS) | say | edge | dashscope | elevenlabs | gemini | gemini-lite | minimax # voices list [lang] | design "" [--out f.wav] | delete Gemini voice library / voice design # captions [zh|en] [zh-max] [--media-end S] SRT (zh / en / bilingual) + captions.json from the TTS timeline; a caption under 1.8 s is # lengthened into the gap after it, up to the picture's end (--media-end, else index.html's data-duration, media/final.mp4, the timeline) @@ -181,7 +181,7 @@ doctor_downloads() { # what the first runs download on their own and how big it hub=${HF_HUB_CACHE:-${HF_HOME:-$HOME/.cache/huggingface}/hub} snap=$(ls -d "$hub/models--${model//\//--}/snapshots/"*/ 2>/dev/null | head -1 || true) if [ "$(uname)" != Darwin ] || [ "$(uname -m)" != arm64 ]; then - echo "- Qwen3-TTS (bin/vh tts's default voice) runs on Apple Silicon only (mlx-audio): use edge, gemini, dashscope or elevenlabs here" + echo "- Qwen3-TTS (bin/vh tts's default voice) runs on Apple Silicon only (mlx-audio): use edge, gemini, minimax, dashscope or elevenlabs here" else command -v uv >/dev/null 2>&1 && ! uv run -q --no-project --offline --with mlx-audio python -c "" >/dev/null 2>&1 && voice_pkgs=1 if [ -n "$snap" ] && [ -e "$(set -- "$snap"*.safetensors; echo "$1")" ]; then echo "$c_ok Qwen3-TTS model $model · $(size_of "$snap")" diff --git a/playbook/04-audio.md b/playbook/04-audio.md index 487f12b..6135edc 100644 --- a/playbook/04-audio.md +++ b/playbook/04-audio.md @@ -13,6 +13,7 @@ bin/vh tts projects/

qwen Serena zh # 本地开源 Qwen3-TTS(默认;首次下载约 2GB) bin/vh tts projects/

qwen Aiden en # 同一份稿子,英文旁白 bin/vh tts projects/

gemini Kore zh --align gemini # 云端 Gemini;--align:逐句词级时间 + ASR 对稿,跑偏的句子会被标出来 +bin/vh tts projects/

minimax "Chinese (Mandarin)_Reliable_Executive" zh --join all # 云端 MiniMax:整份稿一个请求,音色不漂,自带逐字时间 bin/vh captions projects/

zh # → captions.zh/en/bi.srt + captions.json(给引擎画进画面) bin/vh voices list en-GB --gender female # Gemini 音色库;bin/vh voices design "<描述>" 设计一个音色 → voice_… id # 配乐:代码作曲,段落对齐镜头,节拍精确 @@ -39,7 +40,7 @@ bin/vh beats <任意音乐文件> # 外来音乐的节 - 双语写成 `中文 || English`,`--lang` 决定念哪一边,两边都会进时间表; - 只写一边的句子照原样念。进字幕时,含中日韩字符的算中文,不含的算英文,所以纯英文稿出的是 `captions.en.srt`。想让一句纯英文(比如品牌名)留在中文一侧,写成 `Claude Code ||`;只有标点的一句(`……`)跟着前面的单边句走; - 可以加 `@id` 前缀,比如 `@hook 一句话,做出一支片子。 || One sentence in, one film out.`。 -- **写念法,不写字面**:单位、缩写和符号,TTS 常照字母或字面念。showcase 02 的"2 GHz"被念成了 G、H、Z 三个字母,维护者听出来才重录。送去合成的文字要写真实的念法("2 G赫兹""50 千赫兹"),数字、公式、英文缩写也一样。 +- **写念法,不写字面**:单位、缩写和符号,TTS 常照字母或字面念。showcase 02 的"2 GHz"被念成了 G、H、Z 三个字母,维护者听出来才重录。送去合成的文字要写真实的念法("2 G赫兹""50 千赫兹"),数字、公式、英文缩写也一样。年份也是:MiniMax 把"1953 年"念成"一千九百五十三年",要写成"一九五三年"。 - `script.txt` 一句只有一份文本,合成和字幕都用它。念法和字幕必须不同(比如画面烧着"2 GHz")时,就用念法单独合成那一句,timeline 里的字幕文字留原文,再在 `script.txt` 的注释里记下念法(02 就是这样做的)。 - 合成后,单位和数字要人耳听一遍:Gemini 转写的第一遍对这类词也不可靠(02 的新 take 第一遍只有 0.10),不能只看相似度。 - 每句合成后,provider 自带的句首句尾静音会被裁掉,所以句间距离就是 `--gap`,按拍落点时人声也落在拍上(裁法和阈值见 `tools/audio/README.md`;想保留原样,加 `--keep-edges`); @@ -59,6 +60,7 @@ bin/vh beats <任意音乐文件> # 外来音乐的节 |---|---| | `qwen`(默认) | 本地开源 Qwen3-TTS,免费、离线,首次下载约 2 GB;0.6B 模型念英文短句偶尔停不下来,每条都要查 | | `gemini` | 云端,表演力最好,能用一句话导演语气,适合讲解旁白和双人对话;要 `GEMINI_API_KEY` | +| `minimax` | 云端 MiniMax(`speech-2.8-hd`),中文音色多(`Chinese (Mandarin)_*`)。`--join all` 整份稿一个请求,从头到尾一个声音;自带逐字时间,不用转写;支持停顿标记 `<#0.4#>` 和发音词典 `@pronounce`。语气只认一个情绪词。要 `MINIMAX_TOKEN_PLAN_KEY` 或 `MINIMAX_API_KEY`,细节见 `tools/audio/README.md` 的"MiniMax" | | `say` / `edge` | 打草稿;只能念,不能演 | | `elevenlabs` / `dashscope` | 接口已留,还没实测 | @@ -184,10 +186,11 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词 - 写情绪、语速、重音、句尾走向、停顿,例如"像在跟朋友分享一个惊人的发现,语速偏快,'一句话'重读,尾音上扬成问句"; - 用"像在……"打比方,比"专业""自然"这种抽象词好用得多; - 同一段里相邻两句的语气要有落差:问句接答句,铺垫接爆点,快接慢。 -- **标签**:``、``、``、``、`` 可以直接写进句子里。只有 `gemini` 会演出来,其他 provider 和字幕都会自动去掉(规则见 `tools/audio/README.md`)。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。 +- **标签**:``、``、``、``、`` 可以直接写进句子里。只有 `gemini` 会演出来,其他 provider 和字幕都会自动去掉(规则见 `tools/audio/README.md`)。`minimax` 用的是停顿标记 `<#0.4#>`(秒),同样只有它会停。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。 - **给 gemini 的指示要短**:一句话讲清情绪和节奏就够了。年龄、性别、口音属于音色,不要写进指示;长段的人设和导演笔记容易让音色漂移。 - **谁能演**: - `gemini` 表演力最好,整体和逐句指示都听; + - `minimax` 只认一个情绪词(`calm`、`happy`、`whisper` 等),写别的会被忽略; - 本地 `qwen` 的 0.6B 模型不接受指示,要换 1.7B 的 instruct 模型; - `say` 和 `edge` 只能念。 - **写法示例**(逐句导演,后两句用了下文的对拍写法): @@ -272,7 +275,7 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词 | 需求 | 方案 | |---|---| | 中文配音,本地 | `mlx-audio` 在 Apple Silicon 上跑 Qwen3-TTS(Apache-2.0,支持方言和声音设计)。要克隆声音用 CosyVoice 或 GPT-SoVITS。 | -| 配音,云端,追求稳定 | ElevenLabs 的 `/v1/text-to-speech/{voice_id}/with-timestamps` 直接返回字符级时间(`bin/vh tts` 拼成词级);讲解旁白要表演力、或者要双人对话时,用 Gemini 3.8 Flash TTS(`bin/vh tts … gemini`,能用一句话导演语气,有免费档);国内可选火山豆包、阿里百炼、MiniMax。 | +| 配音,云端,追求稳定 | ElevenLabs 的 `/v1/text-to-speech/{voice_id}/with-timestamps` 直接返回字符级时间(`bin/vh tts` 拼成词级);讲解旁白要表演力、或者要双人对话时,用 Gemini 3.8 Flash TTS(`bin/vh tts … gemini`,能用一句话导演语气,有免费档);MiniMax 已封装(`bin/vh tts … minimax`,自带逐字时间,整份稿一个请求音色不漂);国内还可选火山豆包、阿里百炼。 | | 免费、先凑合用 | `edge-tts`(微软中文音色,非官方接口,随时可能失效) | | 词级时间戳 | 已封装:`bin/vh tts … --align gemini`(Gemini 3.5 Transcribe,词级,0.1 s 步长,顺带对稿)。离线的话,中文用 FunASR(字级,带标点);通用用 whisper.cpp(Mac 上有 Metal 加速);已有讲稿或歌词、只需对齐时用 ctc-forced-aligner | | 节拍 | librosa 或 beat_this(后者 downbeat 更准)。音乐平缓时,检测出来的 BPM 只是一个强加的节拍器,不能拿来硬切。 | @@ -301,7 +304,7 @@ bin/vh qa mix $A/stems --beats $A/music.beats.json # 只看混音报告(词 ``` - `cues` 对应 `SCRIPT.md` 里的 `{cue}` 标记,用来在某个词出现的那一刻触发画面。 -- `words` 由 `bin/vh tts … --align gemini`(或 elevenlabs)写入;对稿结果 `asr` 和对话里的 `speaker` 也在每个 segment 里。 +- `words` 由 `bin/vh tts … --align gemini`(或 elevenlabs、minimax)写入;对稿结果 `asr` 和对话里的 `speaker` 也在每个 segment 里。 - ClaudeAnimationBase 用 `src/config.js` 的 `bpm` 和 `offset`,`pulse()` 和 `beatN()` 会自动对齐节拍。 - Remotion 的做法是在 `calculateMetadata` 里读取音频时长来设置帧数;HyperFrames 的做法是 `sync-durations` 用实测时长覆盖估计值。 diff --git a/tools/audio/README.md b/tools/audio/README.md index 6ebdcea..1583606 100644 --- a/tools/audio/README.md +++ b/tools/audio/README.md @@ -18,6 +18,7 @@ | `dashscope` | 阿里云百炼 Qwen3-TTS(`qwen3-tts-flash`),云端 | `DASHSCOPE_API_KEY` | 句级 | 接口已留,待实测 | | `elevenlabs` | 高质量多语种配音 | `ELEVENLABS_API_KEY`、voice_id | **词级**(接口给字符级,`bin/vh tts` 拼成词) | 接口已留,待实测 | | `gemini` / `gemini-lite` | Google Gemini 3.8 Flash TTS(2026-09-23 发布)/ Flash-Lite TTS,云端。表演力强,适合讲解旁白和双人对话。Flash 支持 130 种语言,Lite 支持 101 种,都含中文,语言按文本自动识别。`--voice` 用 30 个 studio 音色之一(默认 Kore,每个都能说所有支持的语言)、音色库 id 或自己设计的 `voice_…`(见下文"音色");`bin/vh tts` 的第 5 个参数(`--instruct`)用一句话描述语气,例如"平静、笃定的纪录片旁白";稿子里可以直接写 ``、``、`` 这类标签,只有 gemini 会演出来,其他 provider 和字幕都会自动去掉 | `GEMINI_API_KEY`(Google AI Studio)。有免费档,但免费档的内容可能被 Google 用来改进产品;付费价格见下文"费用" | 句级;加 `--align gemini` 后为词级 | ✅ `gemini`:2026-09-30 实测中英(Kore / Charon),一次通过,逗号处有自然停顿;双人对话、整段合成、设计音色同日实测。`gemini-lite` 走同一条代码路径(只换模型 id),没有单独实测 | +| `minimax` | MiniMax T2A v2(`speech-2.8-hd`),国际站 `api.minimax.io`,云端。中文系统音色是 `Chinese (Mandarin)_*`(默认 `Chinese (Mandarin)_Reliable_Executive`),英文默认 `English_expressive_narrator`。句子里可以写停顿标记 `<#0.4#>`,稿子里可以写发音词典 `@pronounce`;`--join all` 整份稿一个请求,音色不漂。见下文"MiniMax" | `MINIMAX_TOKEN_PLAN_KEY`(Token Plan,优先)或 `MINIMAX_API_KEY`(按量付费) | **词级**(请求自带逐字时间,不用转写) | ✅ 2026-10-07 实测中英:逐句、`--join block`、`--join all`、`--resume` | **已知问题**:默认的本地 Qwen3-TTS 0.6B 模型念英文短句时偶尔停不下来。实测 "One sentence in, a film out." 念完后又多出约 10 秒低电平的含糊声音,整句 12.6 s。所以每条配音都要查一遍:加 `--align gemini` 让机器查(见下一节),或者至少看一眼 `timeline..json` 里的时长。被标出来的那句重跑,或者换 1.7B 模型、换 `gemini`。同一稿 Gemini 的四条都正常。 @@ -70,7 +71,7 @@ ### 整段合成(`--join block|all`) -`--join block` 把空行之间的连续几句放进一个请求,`--join all` 把整份稿放进一个请求(gemini 单个请求最多 8,192 个输入 token)。句与句之间的语气是连着的,比一句一句拼起来自然。行边界从词级时间反推,所以会自动打开 `--align gemini`,`vo//NN.wav` 是从整段音频里切出来的。其他 provider 也能用,只是把几句拼成一段文字去念。 +`--join block` 把空行之间的连续几句放进一个请求,`--join all` 把整份稿放进一个请求(gemini 单个请求最多 8,192 个输入 token)。句与句之间的语气是连着的,比一句一句拼起来自然。行边界从词级时间反推,所以会自动打开 `--align gemini`,`vo//NN.wav` 是从整段音频里切出来的。其他 provider 也能用,只是把几句拼成一段文字去念。`minimax` 例外:句间插 `--gap` 秒的停顿标记,行边界用它自己返回的逐字时间,不用转写(见下文"MiniMax")。 代价是没法逐句控制: - 某一句念得不好,只能整段重跑; @@ -98,8 +99,37 @@ - **语气写短**:`--instruct` 和 `[指示]` 最后都进 `speech_metadata.style`。官方建议只写这一句的情境语气("压低声音,卖个关子");年龄、性别、口音这类身份特征不要写进 style,要换音色,或者用 `voices design` 做一个。长篇的"角色设定""导演笔记"是音色漂移最常见的原因; - **每段 Gemini 音频都带 SynthID 水印**(听不出来,但能检测到)。用了 Gemini 旁白的片子,在项目 `NOTES.md` 里写明"旁白为 AI 合成(Gemini TTS,含 SynthID 水印)",发布时按平台要求标注。 +### MiniMax(`minimax`) + +```bash +bin/vh tts projects/

minimax "Chinese (Mandarin)_Reliable_Executive" zh --join all # 整份稿一个请求 +MINIMAX_TTS_SPEED=1.25 bin/vh tts projects/

minimax - zh --join all --gap 0.3 # 默认音色,语速 1.25,句间停 0.3 s +``` + +- **key 和接口**:请求发到国际站 `https://api.minimax.io/v1/t2a_v2`。key 只从环境变量读,先找 `MINIMAX_TOKEN_PLAN_KEY`(Token Plan 的 Subscription Key,花的是买的 Credits),没有再用 `MINIMAX_API_KEY`(按量付费,账户没有余额时报 `1008 insufficient balance`)。买了 Credits 却只设了按量付费的 key,也是 1008。国际站的 key 发到国内站 `api.minimaxi.com` 会报 `2049 invalid api key`,所以默认走国际站;用国内站(platform.minimaxi.com)的 key 时,设 `MINIMAX_BASE_URL=https://api.minimaxi.com`。报错只写 key 来自哪个变量,不写 key 本身; +- **音色**:`--voice` 写系统音色的 id。中文音色都以 `Chinese (Mandarin)_` 开头,例如 `Chinese (Mandarin)_Reliable_Executive`(沉稳的中年男声,默认)、`_News_Anchor`(女声新闻主播)、`_Male_Announcer`、`_Radio_Host`;英文默认 `English_expressive_narrator`。全部系统音色(2026-10-07 有 332 个)用 `POST /v1/get_voice`、body `{"voice_type": "system"}` 列出,`bin/vh voices` 只管 Gemini: + + ```bash + curl -s https://api.minimax.io/v1/get_voice -H "Authorization: Bearer $MINIMAX_TOKEN_PLAN_KEY" -H "Content-Type: application/json" -d '{"voice_type":"system"}' | python3 -c 'import json,sys; [print(v["voice_id"], "|", v.get("voice_name", "")) for v in json.load(sys.stdin)["system_voice"]]' + ``` + + id 里有空格,命令行里要加引号;也因为有空格,写不进 `@speakers`,双人对话暂时用别家; +- **参数**:`MINIMAX_TTS_MODEL`(默认 `speech-2.8-hd`)、`MINIMAX_TTS_SPEED`(语速 0.5–2,默认 1)、`MINIMAX_TTS_VOL`(音量 0–10,默认 1)、`MINIMAX_TTS_PITCH`(音调 −12–12,默认 0)。超出范围在合成前就报错。这几个值记在 `vo//_run.json` 里,`--resume` 时和那次不一样就拒绝。`language_boost` 按 `--lang` 给 `Chinese` 或 `English`,`text_normalization` 打开(数字、符号按念法读,见下面的"年份"); +- **语气**:MiniMax 不听一句话的导演指示,只认一个情绪词:`happy`、`sad`、`angry`、`fearful`、`disgusted`、`surprised`、`calm`、`fluent`、`whisper`。`--instruct` 和 `[ ]` 里写这些词才有用(两处都写了情绪词时取 `[ ]` 的),别的文字会被忽略,命令提示一次。整段合成时只用 `--instruct`,逐句的 `[ ]` 不进请求; +- **词级时间不用转写**:请求时打开字幕(`subtitle_type: word`),MiniMax 返回每个字的起止(毫秒)。`bin/vh tts` 把它们拼回稿子的写法:中文一字一个 word,英文一词一个,"1953" 这样的数字是一个 word,覆盖它念出来的全部音节,`@pronounce` 里的词也一样。所以 timeline 里直接有 `words`,`bin/vh captions` 可以逐字出字幕;`--join` 也不会因此打开转写,不要 `GEMINI_API_KEY`。想要 ASR 对稿时照样可以加 `--align gemini`:整段合成时每段转写一次,相似度按句算,句子的起止仍用 MiniMax 的时间; +- **整份稿一个请求,音色不漂**:逐句或逐段分开合成时,段与段之间能听出像换了个人。`--join all` 把整份稿放进一个请求(不超过 10,000 字符,中文旁白约半小时),从头到尾是同一口气,讲解、纪录片这类旁白推荐这样做。句与句之间插一个 `--gap` 秒的停顿标记(默认 0.25 s),再按 MiniMax 的逐字时间把每句切回 `vo//NN.wav`。代价和"整段合成"一节一样:不能单独重念一句,也不能逐句卡拍。`--resume` 时,整段音频和它的逐字时间(`_blockNN.words.json`)都在,才算这一段做完; +- **停顿标记**:句子里写 `<#0.4#>`(秒,0.01–99.99),MiniMax 在那里停 0.4 s;字幕和其他 provider 都会去掉它。写在一句末尾的标记,整段合成时加到这一句和下一句之间的停顿上,比如一段讲完想多停一会儿,就在这段最后一句末尾写 `<#0.5#>`。MiniMax 要求标记夹在两段要念的文字之间,所以一个请求开头和结尾的标记(逐句合成时就是句首和句尾)会被丢掉;连着写的几个标记合成一个; +- **发音词典**:人名、多音字念错时,在 `script.txt` 里加一行 `@pronounce 硖合/(xia2)(he2)`(拼音加声调数字,一个字一组括号;英文词可以写 `omg/oh my god`),可以写多行,不带空格的几条也可以写在同一行。它们进请求的 `pronunciation_dict.tone`,只对 `minimax` 生效(gemini 没有发音词典,要改写成同音字)。改了词典,`--resume` 会拒绝,要重新合成; +- **年份会被念成数**:开着 `text_normalization`,"1953 年"念成"一千九百五十三年"。送去合成的文字里,年份写成"一九五三年"。2001、2003 念成"两千零一""两千零三",可以接受。`script.txt` 的文字同时是字幕,字幕要显示"1953 年"时,按 `playbook/04-audio.md`"写念法"的办法单独处理那一句; +- **计费**:按字符算(响应里的 `usage_characters`)。实测一个汉字算 2 个,字母、数字、标点、空格和停顿标记各算 1 个; +- **实测**(2026-10-07,`speech-2.8-hd`): + - 三句中文 `--join all`:一个请求,整条命令约 9 s,出 10.3 s 音频,三句的起止和逐字时间都来自字幕;句中 `<#0.4#>` 处实际停了约 0.8 s(加上逗号自己的停顿); + - 逐句合成时,每句前后各有约 0.3 s 空白,会被自动裁掉; + - 一支片子里 `Chinese (Mandarin)_Reliable_Executive` 在语速 1.25 下约 5.35 字/秒(说话时)。 + ### 旁白里的标签 +- **停顿标记**:`<#0.4#>` 是 MiniMax 的停顿(秒),只有 `minimax` 会停,其他 provider 和字幕都会去掉它,见上一节。 - **标签**:``、``、``、``、`` 可以直接写进句子里。只有 `gemini` 会演出来;其他 provider 和字幕都会自动去掉。去掉的规则:这 5 个和 `` 一律去掉;别的 `<词>` 只要不是两边都紧贴字母或数字,也当标签去掉,所以 `xz` 这类式子会原样保留。中文稿里也写英文标签,官方说这样效果最好。标签只管某一刻的动作(停顿、呼吸、笑),持续的语气写在 `[ ]` 里。 ## 外来音乐的节拍(`beats`) diff --git a/tools/audio/captions.py b/tools/audio/captions.py index cdc55a4..e976d3c 100644 --- a/tools/audio/captions.py +++ b/tools/audio/captions.py @@ -8,7 +8,7 @@ (HyperFrames / p5 / Manim read it; every caption stays a pure function of t). Soft subtitle tracks for players/platforms: bin/vh mux

minimax - zh --join all --gap 0.3 ``` id 里有空格,命令行里要加引号;也因为有空格,写不进 `@speakers`,双人对话暂时用别家; -- **参数**:`MINIMAX_TTS_MODEL`(默认 `speech-2.8-hd`)、`MINIMAX_TTS_SPEED`(语速 0.5–2,默认 1)、`MINIMAX_TTS_VOL`(音量 0–10,默认 1)、`MINIMAX_TTS_PITCH`(音调 −12–12,默认 0)。超出范围在合成前就报错。这几个值记在 `vo//_run.json` 里,`--resume` 时和那次不一样就拒绝。`language_boost` 按 `--lang` 给 `Chinese` 或 `English`,`text_normalization` 打开(数字、符号按念法读,见下面的"年份"); +- **参数**:`MINIMAX_TTS_MODEL`(默认 `speech-2.8-hd`)、`MINIMAX_TTS_SPEED`(语速 0.5–2,默认 1)、`MINIMAX_TTS_VOL`(音量 0.01–10,默认 1)、`MINIMAX_TTS_PITCH`(音调 −12–12,默认 0)。超出范围在合成前就报错。这几个值记在 `vo//_run.json` 里,`--resume` 时和那次不一样就拒绝。`language_boost` 按 `--lang` 给 `Chinese` 或 `English`,`text_normalization` 打开(数字、符号按念法读,见下面的"年份"); - **语气**:MiniMax 不听一句话的导演指示,只认一个情绪词:`happy`、`sad`、`angry`、`fearful`、`disgusted`、`surprised`、`calm`、`fluent`、`whisper`。`--instruct` 和 `[ ]` 里写这些词才有用(两处都写了情绪词时取 `[ ]` 的),别的文字会被忽略,命令提示一次。整段合成时只用 `--instruct`,逐句的 `[ ]` 不进请求; - **词级时间不用转写**:请求时打开字幕(`subtitle_type: word`),MiniMax 返回每个字的起止(毫秒)。`bin/vh tts` 把它们拼回稿子的写法:中文一字一个 word,英文一词一个,"1953" 这样的数字是一个 word,覆盖它念出来的全部音节,`@pronounce` 里的词也一样。所以 timeline 里直接有 `words`,`bin/vh captions` 可以逐字出字幕;`--join` 也不会因此打开转写,不要 `GEMINI_API_KEY`。想要 ASR 对稿时照样可以加 `--align gemini`:整段合成时每段转写一次,相似度按句算,句子的起止仍用 MiniMax 的时间; -- **整份稿一个请求,音色不漂**:逐句或逐段分开合成时,段与段之间能听出像换了个人。`--join all` 把整份稿放进一个请求(不超过 10,000 字符,中文旁白约半小时),从头到尾是同一口气,讲解、纪录片这类旁白推荐这样做。句与句之间插一个 `--gap` 秒的停顿标记(默认 0.25 s),再按 MiniMax 的逐字时间把每句切回 `vo//NN.wav`。代价和"整段合成"一节一样:不能单独重念一句,也不能逐句卡拍。`--resume` 时,整段音频和它的逐字时间(`_blockNN.words.json`)都在,才算这一段做完; -- **停顿标记**:句子里写 `<#0.4#>`(秒,0.01–99.99),MiniMax 在那里停 0.4 s;字幕和其他 provider 都会去掉它。写在一句末尾的标记,整段合成时加到这一句和下一句之间的停顿上,比如一段讲完想多停一会儿,就在这段最后一句末尾写 `<#0.5#>`。MiniMax 要求标记夹在两段要念的文字之间,所以一个请求开头和结尾的标记(逐句合成时就是句首和句尾)会被丢掉;连着写的几个标记合成一个; -- **发音词典**:人名、多音字念错时,在 `script.txt` 里加一行 `@pronounce 硖合/(xia2)(he2)`(拼音加声调数字,一个字一组括号;英文词可以写 `omg/oh my god`),可以写多行,不带空格的几条也可以写在同一行。它们进请求的 `pronunciation_dict.tone`,只对 `minimax` 生效(gemini 没有发音词典,要改写成同音字)。改了词典,`--resume` 会拒绝,要重新合成; +- **整份稿一个请求,音色不漂**:逐句或逐段分开合成时,段与段之间能听出像换了个人。`--join all` 把整份稿放进一个请求(不超过 10,000 字符,中文旁白约半小时),从头到尾是同一口气,讲解、纪录片这类旁白推荐这样做。句与句之间插一个 `--gap` 秒的停顿标记(默认 0.25 s),再按 MiniMax 的逐字时间把每句切回 `vo//NN.wav`。代价和"整段合成"一节一样:不能单独重念一句,也不能逐句卡拍。这时 `--gap` 已经写进了请求,`--resume` 不能改它(别的时间参数照样可以改);整段音频和它的逐字时间(`_blockNN.words.json`)都在,才算这一段做完。一个请求超过 10,000 字符时,命令在删掉上一次的配音之前就报错; +- **停顿标记**:句子里写 `<#0.4#>`(秒,0.01–99.99),MiniMax 在那里停 0.4 s;字幕和其他 provider 都会去掉它。写在一句开头或末尾的标记,加到这一句和前后一句之间的停顿上:同一个请求里加进连接两句的停顿标记,分开的请求之间加进静音,逐句、分段、整段合成都一样。比如一段讲完想多停一会儿,就在这段最后一句末尾写 `<#0.5#>`。第一句前面和最后一句后面的标记会被丢掉;连着写的几个标记合成一个(MiniMax 不接受连续的标记); +- **发音词典**:人名、多音字念错时,在 `script.txt` 里加一行 `@pronounce 硖合/(xia2)(he2)`(拼音加声调数字,一个字一组括号;英文词可以写 `omg/oh my god`),可以写多行,不带空格的几条也可以写在同一行。没有 `/` 的 `@pronounce …` 仍是一句 id 为 pronounce 的旁白。它们进请求的 `pronunciation_dict.tone`,只对 `minimax` 生效(gemini 没有发音词典,要改写成同音字)。改了词典,`--resume` 会拒绝,要重新合成; - **年份会被念成数**:开着 `text_normalization`,"1953 年"念成"一千九百五十三年"。送去合成的文字里,年份写成"一九五三年"。2001、2003 念成"两千零一""两千零三",可以接受。`script.txt` 的文字同时是字幕,字幕要显示"1953 年"时,按 `playbook/04-audio.md`"写念法"的办法单独处理那一句; - **计费**:按字符算(响应里的 `usage_characters`)。实测一个汉字算 2 个,字母、数字、标点、空格和停顿标记各算 1 个; - **实测**(2026-10-07,`speech-2.8-hd`): diff --git a/tools/audio/tts.py b/tools/audio/tts.py index 08e35e5..ed172ad 100644 --- a/tools/audio/tts.py +++ b/tools/audio/tts.py @@ -107,15 +107,16 @@ MINIMAX_BASE_URL=https://api.minimaxi.com. --voice: a system voice id (default zh "Chinese (Mandarin)_Reliable_Executive", en "English_expressive_narrator"; the Chinese ones are "Chinese (Mandarin)_*", from POST /v1/get_voice {"voice_type": "system"}). - MINIMAX_TTS_MODEL (default speech-2.8-hd), MINIMAX_TTS_SPEED (0.5–2, default 1), MINIMAX_TTS_VOL (0–10, + MINIMAX_TTS_MODEL (default speech-2.8-hd), MINIMAX_TTS_SPEED (0.5–2, default 1), MINIMAX_TTS_VOL (0.01–10, default 1), MINIMAX_TTS_PITCH (−12–12, default 0). --instruct / [direction]: used only when it is one of MiniMax's emotions (happy sad angry fearful disgusted surprised calm fluent whisper). Word timing comes from the request's own subtitles (per character, ms), no ASR needed: --join block|all places the lines of a joined request by them, and --align gemini is off unless asked for. - Pause marks <#0.4#> (seconds) inside a line are voiced; captions and the other providers drop them. In a - joined request the lines are linked by a <#--gap#> mark, plus any mark written at a line's edge; a mark - at the edge of a request is dropped (MiniMax wants one between two spoken words). --join all sends the - whole script in one request (under 10,000 characters), so the voice cannot drift between blocks. + Pause marks <#0.4#> (seconds) inside a line are voiced; captions and the other providers drop them. A mark + at a line's edge adds to the pause between it and the line next to it (inside a joined request: the + <#--gap#> mark that links them; between requests: the silence); one before the first line or after the + last is dropped. --join all sends the whole script in one request (under 10,000 characters), so the + voice cannot drift between blocks; --gap is then part of the text, so --resume keeps it. "@pronounce 硖合/(xia2)(he2)" lines in script.txt (one entry each, or several without spaces) become pronunciation_dict.tone. text_normalization is on: "1953 年" is read 一千九百五十三年, so write years as 一九五三年 in the spoken text. @@ -447,7 +448,7 @@ def p_minimax(text, voice, out: Path, tmp: Path, lang, instruct=None): voice = voice or ("Chinese (Mandarin)_Reliable_Executive" if lang.startswith("zh") else "English_expressive_narrator") audio, subs = minimax_tts(text, voice, lang, minimax_emotion(instruct)) raw = tmp / "mm.wav"; raw.write_bytes(audio); to_wav(raw, out) - words = minimax_words(subs, strip_tags(text)) + words = minimax_words(subs, strip_tags(text, reactions=False)) # prepare() already dropped backchannels if words is None: raise MinimaxError(f"minimax: the subtitle file does not match the text sent ({json.dumps(subs, ensure_ascii=False)[:300]})") return words @@ -468,10 +469,14 @@ def _tag(m): glued = a > 0 and s[a - 1].isascii() and s[a - 1].isalnum() and b < len(s) and s[b].isascii() and s[b].isalnum() return m.group() if glued and m.group(1).lower() not in TAGS else " " -def _pause(m): - """A dropped pause mark leaves a space where the text had one around it, or between two letters/digits.""" +def _spaced(m): + """A run of pause marks stands for a space where the text had one around it, or between two ASCII letters/digits.""" s, a, b = m.string, m.start(), m.end() - return " " if m.group() != m.group().strip() or (a > 0 and b < len(s) and s[a - 1].isalnum() and s[b].isalnum()) else "" + word = lambda c: c.isascii() and c.isalnum() + return m.group() != m.group().strip() or (a > 0 and b < len(s) and word(s[a - 1]) and word(s[b])) + +def _pause(m): + return " " if _spaced(m) else "" def pause_secs(run): return sum(float(x) for x in re.findall(r"\d+(?:\.\d+)?", run)) @@ -491,7 +496,7 @@ def split_pauses(text): m = re.search(PAUSES.pattern + "$", text) if m: trail, text = pause_secs(m.group()), text[:m.start()] - return lead, PAUSES.sub(lambda r: pause_mark(pause_secs(r.group()), r.group() != r.group().strip()), text), trail + return lead, PAUSES.sub(lambda r: pause_mark(pause_secs(r.group()), _spaced(r)), text), trail def strip_tags(text: str, reactions=True, pauses=True) -> str: """Remove inline performance tags, MiniMax pause marks and (in a dialogue script) |reaction| markers: only gemini @@ -521,13 +526,10 @@ def parse_script(path: Path): if head: speakers.update(p.split("=", 1) for p in head.group(1).split()) continue - head = re.fullmatch(r"@pronounce\s+(.+)", line) # @pronounce 硖合/(xia2)(he2) — minimax only - if head: # several on a line when none has a space - items = head.group(1).split() - for x in (items if all("/" in x for x in items) else [head.group(1).strip()]): - if "/" not in x: - sys.exit(f"{path.name}: @pronounce {x!r}: write text/reading, e.g. 硖合/(xia2)(he2)") - pronounce.append(x) + head = re.fullmatch(r"@pronounce\s+(.+/.+)", line) # @pronounce 硖合/(xia2)(he2) — minimax only; without + if head: # a "/" it is a line with the id "pronounce", as before + items = head.group(1).split() # several on a line when none has a space + pronounce += items if all("/" in x for x in items) else [head.group(1).strip()] continue sid = style = spk = None if line.startswith("@") and " " in line: @@ -972,10 +974,30 @@ def main(): print(" --join recovers line boundaries from word timestamps → --align gemini is on") if (gemini or align) and not os.environ.get("GEMINI_API_KEY"): # before anything is synthesized or deleted sys.exit("set GEMINI_API_KEY (Google AI Studio → Get API key)" + ("" if gemini else ": --align gemini / --join transcribe with Gemini")) + if man and a.provider == "minimax" and join != "none" and man.get("gap") != a.gap: + sys.exit(f"--resume: --gap was {man.get('gap')} for the run, now {a.gap}; in a joined minimax request it is part of the" + f" text: pass --gap {man.get('gap')}, or drop --resume to start over") + spoken = lambda s: (s["zh"] if a.lang == "zh" else s["en"]) or s["zh"] or s["en"] + def linked(texts): # one minimax request: the lines linked by a pause mark of --gap s, plus any mark at their edges + text, carry = "", 0.0 + for k, x in enumerate(texts): + lead, core, trail = split_pauses(x) + text += (pause_mark(carry + a.gap + lead, a.lang == "en") if k else "") + core; carry = trail + return text + blocks = [] + for s in segs: + if blocks and join != "none" and (join == "all" or blocks[-1][-1]["_block"] == s["_block"]): + blocks[-1].append(s) + else: + blocks.append([s]) if a.provider == "minimax": minimax_settings() # a value out of range exits here, not mid-run if not minimax_key()[0]: sys.exit("set MINIMAX_TOKEN_PLAN_KEY (Token Plan) or MINIMAX_API_KEY (pay-as-you-go), from platform.minimax.io") + size, first = max(((len(linked([strip_tags(spoken(s), reactions, pauses=False) for s in blk])), blk[0]["id"]) for blk in blocks), default=(0, "")) + if size > MINIMAX_MAX_CHARS: # before the last take is deleted + sys.exit(f"minimax: the request from @{first} would be {size} characters, over the limit of {MINIMAX_MAX_CHARS}: " + + {"all": "use --join block (blank lines in script.txt split it)", "block": "split that block with a blank line"}.get(join, "split the line")) if not a.resume: shutil.rmtree(vo, ignore_errors=True); vo.mkdir(parents=True) grids = {} @@ -994,21 +1016,22 @@ def main(): have = ", ".join(k for k in ("beats", "downbeats") if bm.get(k)) or "neither beats nor downbeats" sys.exit(f"--beats {a.beats}: no '{a.snap}' grid in it (it has {have}); pass --snap downbeat, or a map written by bin/vh music or bin/vh beats") said = set() - def next_start(first, t, which=None): + def next_start(first, t, which=None, extra=0.0): """First point of the line's grid at or after its earliest start (--lead, else t + --min-gap). Past the end of - the grid (the narration outlives the music) the line follows --gap instead, and that is said once per grid.""" - name = which if which in grids else a.snap; g = grids.get(name, grid); earliest = a.lead if first else t + a.min_gap + the grid (the narration outlives the music) the line follows --gap instead, and that is said once per grid. + extra: minimax pause marks at the edges of the two lines, added to the pause between them.""" + name = which if which in grids else a.snap; g = grids.get(name, grid); earliest = a.lead if first else t + a.min_gap + extra x = next((x for x in g if x >= earliest - 1e-6), None) if x is None and a.beats and name not in said: said.add(name) print(f" warning: the {name} grid " + (f"ends at {g[-1]:.2f} s" if g else f"is empty in {a.beats}") + f"; from here lines follow --gap {a.gap}", flush=True) - return x if x is not None else (a.lead if first else t + a.gap) + return x if x is not None else (a.lead if first else t + a.gap + extra) def silence(sec, k): f = vo / f"_gap{k:02d}.wav" run(["ffmpeg", "-v", "error", "-y", "-f", "lavfi", "-i", f"anullsrc=r={SR}:cl=mono", "-t", f"{max(sec, 0.001):.4f}", str(f)]) return f def prepare(s, conv=False): # → text for the provider; the segment keeps caption-clean sides - raw = (s["zh"] if a.lang == "zh" else s["en"]) or s["zh"] or s["en"] + raw = spoken(s) s["text"] = strip_tags(raw, reactions) # what captions show if conv and REACTION.search(raw): s["_alt"] = with_reactions(raw) # what the ASR hears in that turn @@ -1021,12 +1044,6 @@ def style_of(s): err = lambda e: str(e) if isinstance(e, (GeminiError, MinimaxError)) else f"{type(e).__name__}: {e}" # every line is prepared first, so the run can be recorded (vo//_run.json) before the first provider call: # --resume checks the script against it and continues from whatever files that run left - blocks = [] - for s in segs: - if blocks and join != "none" and (join == "all" or blocks[-1][-1]["_block"] == s["_block"]): - blocks[-1].append(s) - else: - blocks.append([s]) convs = [conversational and any(s.get("speaker") for s in blk) for blk in blocks] # a block without labels stays narration for bi, (blk, conv) in enumerate(zip(blocks, convs)): if conv and not all(s.get("speaker") for s in blk): @@ -1047,7 +1064,8 @@ def style_of(s): else: (vo / "_run.json").write_text(json.dumps({"provider": a.provider, "voice": a.voice, "speakers": speakers, "lang": a.lang, "join": join, "keep_edges": a.keep_edges, "instruct": a.instruct, "style_env": os.environ.get("GEMINI_TTS_STYLE"), - **({"env": penv, "pronounce": PRONOUNCE} if a.provider == "minimax" else {}), + **({"env": penv, "pronounce": PRONOUNCE, **({"gap": a.gap} if join != "none" else {})} + if a.provider == "minimax" else {}), "lines": lines}, ensure_ascii=False, indent=1), encoding="utf-8") t, parts, asr_s, asr_n, trimmed = 0.0, [], 0.0, 0, 0.0 joined = audio / f"voiceover.{a.lang}.wav" @@ -1089,12 +1107,9 @@ def asr_recheck(wav, vocab): # the text-only re-check, cached the same way (sa return txt def shifted(words, h): # word times after trim_edges cut h s from the head return [dict(w, start=max(0.0, w["start"] - h), end=max(0.0, w["end"] - h)) for w in words] if words and h else words - def linked(blk): # one minimax request for a block: the lines linked by a pause mark of --gap s, plus any - text, carry = "", 0.0 # mark written at their edges - for k, s in enumerate(blk): - lead, core, trail = split_pauses(s["_ptext"]) - text += (pause_mark(carry + a.gap + lead, a.lang == "en") if k else "") + core; carry = trail - return text + def edge(s): # seconds of minimax pause marks before and after a line (other providers drop them) + lead, _, trail = split_pauses(s["_ptext"]) if a.provider == "minimax" else (0.0, "", 0.0) + return lead, trail if join == "none": for i, s in enumerate(segs): out = vo / f"{i+1:02d}.wav"; words = None; wfile = out.with_suffix(".words.json") @@ -1113,7 +1128,7 @@ def linked(blk): # one minimax request for a block: the lines linked by a elif timed: words = json.loads(wfile.read_text(encoding="utf-8")) d = duration(out); s["_cutlen"] = d - start = next_start(i == 0, t, s.get("snap")) + start = next_start(i == 0, t, s.get("snap"), edge(segs[i - 1])[1] + edge(s)[0] if i else 0.0) if start > t: parts.append(silence(start - t, i)) t = start @@ -1152,7 +1167,7 @@ def linked(blk): # one minimax request for a block: the lines linked by a turns = [{"text": s["_ptext"], "style": style_of(s) or env_style, "speaker": s.get("speaker") if conv else None} for s in blk] gemini_tts(turns, bout, tmp, gemini_model(a.provider), dict(speakers) if conv else a.voice) elif timed: - pw = PROVIDERS[a.provider](linked(blk), a.voice, bout, tmp, a.lang, a.instruct) + pw = PROVIDERS[a.provider](linked([s["_ptext"] for s in blk]), a.voice, bout, tmp, a.lang, a.instruct) else: PROVIDERS[a.provider](("" if a.lang == "zh" else " ").join(s["_ptext"] for s in blk), a.voice, bout, tmp, a.lang, a.instruct) if not a.keep_edges: # only the block's own edges: pauses inside stay @@ -1173,7 +1188,7 @@ def linked(blk): # one minimax request for a block: the lines linked by a except Exception as e: # without the provider's timing the block stays whole and fail = err(e) # its lines share its span; flagged either way asr_s += time.time() - t0 - start = next_start(bi == 0, t, blk[0].get("snap")) + start = next_start(bi == 0, t, blk[0].get("snap"), edge(blocks[bi - 1][-1])[1] + edge(blk[0])[0] if bi else 0.0) if start > t: parts.append(silence(start - t, bi)) t = start