Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 15 additions & 5 deletions bazel/python/deps.bzl
Original file line number Diff line number Diff line change
Expand Up @@ -14,13 +14,15 @@ entirely in `pyproject.toml`. This is a list of *which marker-
conditional names rules_python skips*, derived from grepping
requirements.lock.txt for `; python_full_version` / `; sys_platform`
markers that don't satisfy our pinned 3.12 + cross-platform set.
Korean G2P's two platform-specific analyzer roots also need an explicit
select because the hub omits them from all_requirements on matching hosts.

Track upstream: bazel-contrib/rules_python#2244 (and friends) — once
fixed, this whole file collapses to `_RUNTIME_DEPS = all_requirements`
directly in the consumer BUILD.
"""

load("@pypi//:requirements.bzl", _all_requirements = "all_requirements")
load("@pypi//:requirements.bzl", _all_requirements = "all_requirements", _requirement = "requirement")

# Names that appear in `all_requirements` but whose BUILD file pip.parse
# elides because of a `python_full_version` / `sys_platform` marker.
Expand Down Expand Up @@ -64,6 +66,9 @@ _MARKER_FILTERED = [
# sys_platform == 'win32'
"pywin32_ctypes",
"tzdata",
# Korean G2P's platform-specific roots are selected explicitly below.
"eunjeon",
"python_mecab_ko",
]

def _is_filtered(label):
Expand All @@ -75,8 +80,13 @@ def _is_filtered(label):
def all_runtime_deps():
"""Every dep the lockfile resolves for the current platform.

No name list maintained anywhere in BUILD/justfile/MODULE — call this
from py_library/py_binary/py_test `deps =` and the dependency set is
implicit in pyproject.toml + requirements.lock.txt.
Call this from py_library/py_binary/py_test `deps =`. Packages come from
pyproject.toml + requirements.lock.txt, with the Korean analyzer selected
for the target platform below.
"""
return [d for d in _all_requirements if not _is_filtered(d)]
# The hub omits these marker-conditional roots from all_requirements even
# on matching hosts. Keep them available to g2pk2 without runtime pip calls.
return [d for d in _all_requirements if not _is_filtered(d)] + select({
"@platforms//os:windows": [_requirement("eunjeon")],
"//conditions:default": [_requirement("python-mecab-ko")],
})
78 changes: 78 additions & 0 deletions book/src/batchalign/user-guide/cli-reference.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,7 @@ or over their input as described below. See [Command I/O](../reference/command-i
|---|---|
| `transcribe` | Recording to CHAT transcript |
| `align` | Forced alignment of CHAT against audio |
| `phonetic` | Add observed IPA to `%pho` using audio and phone-sequence DP |
| `morphotag` | Add `%mor` and `%gra` |
| `utseg` | Revise utterance segmentation |
| `translate` | Add translation tiers |
Expand Down Expand Up @@ -143,6 +144,83 @@ Accepts the shared input selection and `-o/--out` options above.
| `--engine` | pyannote-ai | Choices: `pyannote-ai`, `pyannote`. Diarization engine: pyannote-ai (cloud) or pyannote (local). |
| `--num-speakers`, `-n` | 0 | Expected speaker count; zero auto-detects. |

## phonetic

Accepts timed CHAT and matching audio, using shared input selection and
`-o/--out`. Install the `phonetic` extra for PhoneticXeus, Piper Plus G2P, and Epitran.
Phonetic transcription requires Python 3.11 or newer.
The first inference downloads the pinned model revision;
building the CLI and displaying help do not download model weights.

```bash
just batchalign cli phonetic recording.cha --out phonetic-output --force-cpu
```

The command recognizes phones from audio and DP-aligns them against reference
IPA pronunciations to recover word boundaries. `%pho` retains the observed IPA,
including pronunciation differences. Existing `%pho` tiers are preserved.
The input must have utterance timing bullets; use `utr` first when needed.
Word-level forced alignment is not required.

Like Whisper forced alignment, phonetic inference groups consecutive utterances
into approximately 20-second audio windows, then projects phones back to their
original words and utterances. Gaps over two seconds and backwards timings start
a new window; an utterance longer than 20 seconds stays whole. Words receiving
no phones are marked `…` (uncoded) and reported with their utterance timing.

Windows of similar duration are batched using padding and real encoder lengths.
Normalization is performed independently per window, and padded output frames
are excluded from decoding. CPU defaults to one window at a time; CUDA defaults
to two. Increase `--batch-size` to try higher GPU throughput, or decrease it to
reduce memory use. Batched and single-window output can differ slightly because
the upstream encoder's convolution branches remain sensitive to padding.

| Option | Default | Details |
|---|---|---|
| `--pronunciations` | None | UTF-8 CSV with header `word,ipa`; one word or whole phonological unit and its IPA per row, overriding the generated pronunciation. |
| `--force-cpu` | False | Use CPU instead of automatic CUDA selection. MPS is not selected. |
| `--batch-size` | CPU 1, CUDA 2 | Maximum audio windows per model batch; must be positive. |

The task runner passes the primary `@Languages` code from CHAT. No separate
language option is needed. [Piper Plus G2P](https://pypi.org/project/piper-plus-g2p/)
0.2.0 handles English, Japanese, Mandarin Chinese, Korean, Spanish, French,
Portuguese, and Swedish. Both `cmn` and `zho` select Mandarin; Cantonese (`yue`)
is a separate language and uses the fallback.

Other languages use [Epitran](https://github.com/dmort27/epitran), with their
default script resolved automatically (e.g. Russian `rus-Cyrl`, Hindi
`hin-Deva`). A failure in a supported Piper backend is reported rather than
silently switching providers. English does not require Flite's `lex_lookup`
or eSpeak; Python language packages are included in the extra. Language
resources may download on first use.

The DP compares IPA directly, preserving distinctions such as nasalization,
vowel length, and tone. Piper's language-specific phone/tone labels are
converted to IPA before comparison. References only determine grouping;
they never replace the observed phones. Alternate scripts and code-switched
words can use pronunciation overrides.
Pass `--pronunciations pronunciations.csv` to supply dialect forms or words
in unsupported languages. For example:

```csv
word,ipa
wug,wʌɡ
bonjour,bɔ̃ʒuʁ
the cat,ðəkæt
```

The header is required. Use standard CSV quoting for cells containing commas.
Words are matched without case; empty cells and duplicate words are errors.
An unknown pronunciation, invalid
audio window, or alignment leaving a word without phones fails the file without
overwriting it. Insertions between word anchors attach to the preceding word;
review inferred boundaries, especially around reduced or atypical speech.

Python pipelines can compose `recipes.phonetic(phonetic_backend=backend,
utr_backend=...)` with existing tasks. Custom backends implement the `Phonetic`
marker and typed `PhoneticInput`/`PhoneticOutput` contract; the Rust runner owns
CHAT extraction, result validation, and tier insertion.

## ai

Accepts the shared input selection and `-o/--out` options above.
Expand Down
11 changes: 10 additions & 1 deletion crates/batchalign/batchalign-core/src/base.rs
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ use crate::proto::convert::{ConvertInput, MediaOutput};
use crate::proto::coref::{CorefInput, CorefOutput};
use crate::proto::fa::{FaInput, FaOutput};
use crate::proto::morphosyntax::{MorphosyntaxInput, MorphosyntaxOutput};
use crate::proto::phonetic::{PhoneticInput, PhoneticOutput};
use crate::proto::speaker::{SpeakerInput, SpeakerOutput};
use crate::proto::translate::{TranslateInput, TranslateOutput};
use crate::proto::utr::{UtrInput, UtrOutput};
Expand Down Expand Up @@ -69,6 +70,8 @@ pub enum Task {
Compare,
/// Decode media and encode a new WAV or MP3 artifact.
Convert,
/// Acoustic phonetic transcription into `%pho`.
Phonetic,
}

impl Task {
Expand All @@ -91,6 +94,7 @@ impl Task {
// already get bullets from UtSeg.)
Task::Fa => &[Task::UtSeg, Task::Utr],
Task::Morphosyntax => &[Task::UtSeg],
Task::Phonetic => &[Task::UtSeg, Task::Utr, Task::Fa],
Task::Coref => &[Task::Morphosyntax],
Task::Translate => &[Task::Morphosyntax],
Task::Compare => &[],
Expand All @@ -109,14 +113,15 @@ impl Task {
Task::Utr => "utr",
Task::Morphosyntax => "morphosyntax",
Task::Translate => "translate",
Task::Phonetic => "phonetic",
Task::Coref => "coref",
Task::Compare => "compare",
Task::Convert => "convert",
}
}

/// Every variant — useful for iteration in tests and codegen.
pub const ALL: [Task; 11] = [
pub const ALL: [Task; 12] = [
Task::Ai,
Task::Asr,
Task::Fa,
Expand All @@ -125,6 +130,7 @@ impl Task {
Task::Utr,
Task::Morphosyntax,
Task::Translate,
Task::Phonetic,
Task::Coref,
Task::Compare,
Task::Convert,
Expand Down Expand Up @@ -200,6 +206,7 @@ union_input_output! {
Ai(AiInput) => Ai,
Asr(AsrInput) => Asr,
Fa(FaInput) => Fa,
Phonetic(PhoneticInput) => Phonetic,
Speaker(SpeakerInput) => Speaker,
UtSeg(UtSegInput) => UtSeg,
// UTR's payload is serde-transparent over `AsrInput`, so the
Expand All @@ -217,6 +224,7 @@ union_input_output! {
Ai(AiOutput),
Asr(AsrOutput),
Fa(FaOutput),
Phonetic(PhoneticOutput),
Speaker(SpeakerOutput),
UtSeg(UtSegOutput),
Utr(UtrOutput),
Expand Down Expand Up @@ -252,6 +260,7 @@ try_from_output! {
Ai(AiOutput),
Asr(AsrOutput),
Fa(FaOutput),
Phonetic(PhoneticOutput),
Speaker(SpeakerOutput),
UtSeg(UtSegOutput),
Utr(UtrOutput),
Expand Down
3 changes: 3 additions & 0 deletions crates/batchalign/batchalign-core/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,9 @@ pub use base::{
pub use cache::CacheKey;
pub use metrics::{MetricsArtifact, MetricsKind, MetricsRow, MetricsTable};
pub use proto::convert::{ConvertInput, MediaFormat, MediaOutput};
pub use proto::phonetic::{
PhoneticInput, PhoneticOutput, PhoneticResult, PhoneticUnit, PhoneticUtterance,
};
pub use utils::{
AiChatInput, AudioError, BAError, BAResult, ChatInput, MediaInput, PairedInput, PreparedAudio,
SourceId, SpeakerLabel, prepare_pcm, prepare_pcm_interleaved,
Expand Down
1 change: 1 addition & 0 deletions crates/batchalign/batchalign-core/src/proto/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,7 @@ pub mod compare;
pub mod convert;
pub mod coref;
pub mod fa;
pub mod phonetic;
pub mod morphosyntax;
pub mod speaker;
pub mod translate;
Expand Down
101 changes: 101 additions & 0 deletions crates/batchalign/batchalign-core/src/proto/phonetic.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
//! Acoustic phonetic transcription. CHAT unit ownership survives inference.

use crate::cache::{CacheKey, hash_serialized};
use crate::utils::{PreparedAudio, SourceId};
use schemars::JsonSchema;
use serde::{Deserialize, Serialize};

#[derive(Clone, Debug, Serialize, Deserialize, JsonSchema)]
pub struct PhoneticUnit {
/// Original spoken text, or a CHAT pause to preserve structurally.
pub text: String,
pub pause: bool,
}

#[derive(Clone, Debug, Serialize, Deserialize, JsonSchema)]
pub struct PhoneticUtterance {
/// Line index in the source AST; results must echo it in order.
pub index: usize,
pub start_ms: u64,
pub end_ms: u64,
pub units: Vec<PhoneticUnit>,
}

#[derive(Clone, Debug, Serialize, Deserialize, JsonSchema)]
pub struct PhoneticInput {
pub source_id: SourceId,
pub audio: PreparedAudio,
pub language: String,
pub utterances: Vec<PhoneticUtterance>,
}

impl CacheKey for PhoneticInput {
fn hash(&self, hasher: &mut blake3::Hasher) {
// Crop bounds and unit ownership affect both inference and projection.
hash_serialized(&(&self.audio, &self.language, &self.utterances), hasher);
}
}

#[derive(Clone, Debug, Serialize, Deserialize, JsonSchema)]
pub struct PhoneticResult {
pub index: usize,
/// One IPA string per input unit; pauses must be echoed unchanged.
pub ipa: Vec<String>,
}

#[derive(Clone, Debug, Serialize, Deserialize, JsonSchema)]
pub struct PhoneticOutput {
pub source_id: SourceId,
pub utterances: Vec<PhoneticResult>,
}

crate::register_proto_schema!(PhoneticUnit);
crate::register_proto_schema!(PhoneticUtterance);
crate::register_proto_schema!(PhoneticInput);
crate::register_proto_schema!(PhoneticResult);
crate::register_proto_schema!(PhoneticOutput);

#[cfg(test)]
mod tests {
use super::*;

fn digest(input: &PhoneticInput) -> blake3::Hash {
let mut hasher = blake3::Hasher::new();
input.hash(&mut hasher);
hasher.finalize()
}

#[test]
fn cache_tracks_audio_windows_and_reference_but_not_source_path() {
let mut input = PhoneticInput {
source_id: SourceId::try_new("a.cha").unwrap(),
audio: PreparedAudio {
pcm_f32le: vec![0; 64],
sample_rate: 16000,
channels: 1,
frame_count: 16,
},
language: "eng".into(),
utterances: vec![PhoneticUtterance {
index: 0,
start_ms: 0,
end_ms: 1,
units: vec![PhoneticUnit {
text: "cat".into(),
pause: false,
}],
}],
};
let original = digest(&input);
input.source_id = SourceId::try_new("b.cha").unwrap();
assert_eq!(digest(&input), original);
input.utterances[0].end_ms = 2;
assert_ne!(digest(&input), original);
input.utterances[0].end_ms = 1;
input.utterances[0].units[0].text = "dog".into();
assert_ne!(digest(&input), original);
input.utterances[0].units[0].text = "cat".into();
input.audio.pcm_f32le[0] = 1;
assert_ne!(digest(&input), original);
}
}
2 changes: 2 additions & 0 deletions crates/batchalign/batchalign-core/src/taskrunners/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ pub mod convert;
pub mod coref;
pub mod fa;
mod media;
pub mod phonetic;
pub mod morphosyntax;
pub mod speaker;
pub mod translate;
Expand All @@ -31,6 +32,7 @@ pub fn canonical(task: Task) -> Box<dyn DynTaskRunner> {
Task::Ai => Box::new(ai::AiTaskRunner::default()),
Task::Asr => Box::new(asr::AsrTaskRunner),
Task::Fa => Box::new(fa::FaTaskRunner),
Task::Phonetic => Box::new(phonetic::PhoneticTaskRunner),
Task::Speaker => Box::new(speaker::SpeakerTaskRunner),
Task::UtSeg => Box::new(utseg::UtSegTaskRunner),
Task::Morphosyntax => Box::new(morphosyntax::MorphosyntaxTaskRunner),
Expand Down
Loading
Loading