Open-source implementations for multilingual speech representation, syllabification, and text-independent phonetic alignment
This repository contains the research implementations behind two publications developed during R&D on AI-based pronunciation analysis:
- TIPAA-SSL — Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer
- MUST&P-SRL — Multi-lingual and Unified Syllabification in Text and Phonetic Domains for Speech Representation Learning
The code combines self-supervised speech representations, phonetic processing, frame-level phoneme classification, forced alignment, dynamic time warping, and multilingual linguistic processing.
The original research was motivated by pronunciation analysis and language-learning applications, but the techniques are applicable more broadly to speech analysis, phonetic alignment, pronunciation assessment, speech representation learning, and multilingual speech technology.
Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer
Noé Tits, Prernna Bhatnagar, Thierry Dutoit Acoustics, 2024, 6(3), 772–781.
TIPAA-SSL investigates how the latent representations of a multilingual self-supervised speech model can be used for frame-level phonetic analysis and text-independent phone-to-audio alignment.
Instead of directly using the phoneme sequence predicted by a CTC-based wav2vec 2.0 model, the approach extracts its intermediate/last-layer representations and uses them as a learned phonetic representation.
A lightweight downstream pipeline then performs:
- self-supervised speech representation extraction;
- dimensionality reduction;
- frame-level phoneme classification;
- phoneme probability estimation for each audio frame;
- optional alignment against an expected phoneme sequence using Dynamic Time Warping (DTW).
This provides access to phonetic information while retaining explicit frame-level timing information.
MUST&P-SRL: Multi-lingual and Unified Syllabification in Text and Phonetic Domains for Speech Representation Learning
Noé Tits Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 74–82.
- Paper — ACL Anthology
- DOI
- Implementation
- Main function:
prefill_for_sentence
MUST&P-SRL provides multilingual linguistic processing for extracting structured information from text and phonetic representations.
The approach includes:
- multilingual phonetic transcription;
- syllabification;
- stress information;
- unified representations in text and phonetic domains;
- processing compatible with speech alignment workflows;
- integration with open-source linguistic and speech-processing tools.
The method was evaluated on English, French, and Spanish data and was designed to support speech representation learning and downstream speech-analysis applications.
The two publications address complementary parts of a speech-analysis pipeline.
MUST&P-SRL provides structured linguistic and phonetic representations, including syllable-level information.
TIPAA-SSL provides a way of mapping learned speech representations to frame-level phonetic information and, when an expected phonetic sequence is available, aligning that representation with the audio.
Together, they form part of a broader speech-processing stack for pronunciation analysis and related applications.
The TIPAA-SSL pipeline can be summarized as:
Audio
│
▼
┌────────────────────┐
│ wav2vec 2.0 / │
│ XLS-R │
└─────────┬──────────┘
│
▼
Frame-level latent
representations
│
▼
┌────────────────────┐
│ Dimension │
│ reduction │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ Frame-level │
│ phoneme classifier │
└─────────┬──────────┘
│
▼
Phoneme probability
per frame
│
┌──────┴──────┐
│ │
▼ ▼
Frame-level Expected
phonetic phoneme
analysis sequence
│
▼
Dynamic Time Warping
│
▼
Phone-to-audio alignment
The key idea is to use the learned phonetic space of a multilingual self-supervised model while retaining access to frame-level information instead of relying only on the final collapsed CTC phoneme sequence.
The repository contains the implementation and supporting material for the speech-processing pipeline.
.
├── flowspeech/ # Main speech and linguistic processing package
├── models/ # Trained downstream model components
├── data/ # Data and processing resources
├── scripts/ # Offline processing and experiments
├── syllabipy/ # Modified syllabification component
├── content_tools/ # Linguistic/content-processing interface
├── app/ # API application
├── nginx/ # Deployment configuration
├── notebooks/ # Experimental notebooks
├── README.md # This file
├── INFO.md # Technical implementation notes
├── API_v2.md # API documentation
├── DEPLOYMENT.md # Deployment information
├── INSTALLATION.md # Installation instructions (incl. macOS)
├── MODELS.md # Pretrained model download and setup
├── DATASETS.md # Speech dataset download instructions
├── TESTING.md # How to run the tests
└── DEVELOPMENT.md # Best practices, notebooks and TODO
The most relevant implementation files for the two papers are:
- TIPAA-SSL:
flowspeech/wav2vec2_frame_prediction.py - MUST&P-SRL:
flowspeech/text_processing.py
The project currently targets Python 3.10+.
Some components additionally require:
- Git LFS
- eSpeak / eSpeak NG
- Festival
- micromamba / conda-compatible environment management
The repository contains the downstream model components (dimensionality reduction and frame classifier) in models/. The base wav2vec2 models are downloaded separately, see MODELS.md.
git clone https://github.com/numediart/TIPAA-SSL.git
cd TIPAA-SSLInstall the system dependencies:
sudo apt-get install git-lfs espeak-ng festival
git lfs installThen create the project environment:
micromamba create -n flowspeech_mm -f env.yml
micromamba activate flowspeech_mm
pip install -e .For Apple Silicon / macOS, see the platform-specific instructions in INSTALLATION.md.
See INSTALLATION.md for the full instructions, MODELS.md for the pretrained models, DATASETS.md for the speech datasets, and TESTING.md for how to run the tests.
The pipeline builds on multilingual wav2vec 2.0 / XLS-R representations.
The main pretrained representations used by the research are based on:
facebook/wav2vec2-large-xlsr-53facebook/wav2vec2-xlsr-53-espeak-cv-ft
The latter model was trained for multilingual phoneme recognition. In TIPAA-SSL, its learned representation is reused as a phonetic representation rather than simply taking its final phoneme sequence prediction.
The downstream dimensionality-reduction and frame-classification components used by the pipeline are included in this repository under:
models/
See MODELS.md for how to download the base wav2vec2 models and set up the quantized ONNX model used by the API and tests.
Training the downstream frame-level classifier requires audio frames associated with phoneme labels.
For this purpose, the research pipeline uses forced alignment tools such as the Montreal Forced Aligner (MFA) to obtain phoneme-level timing information.
The resulting aligned data can then be associated with the frame-level representations extracted from the self-supervised speech model.
This makes it possible to train a lightweight classifier on top of the pretrained representation.
See DATASETS.md for instructions on downloading the speech datasets used for training and evaluation.
One important distinction in this work is between:
The model can produce phoneme probability distributions for successive audio frames without requiring a transcript to be provided to the frame classifier.
When an expected phoneme sequence is available, Dynamic Time Warping can be used to align the expected phonemes with the frame-level probability sequence.
This distinction is important for understanding the TIPAA-SSL method and its potential use in pronunciation analysis.
The research was originally developed in the context of AI-based pronunciation analysis and feedback.
A pronunciation-analysis system can use the resulting phonetic information to investigate questions such as:
- Which phoneme was produced?
- When does a phoneme occur in the audio?
- Was an expected phoneme omitted or substituted?
- How closely does the pronunciation follow an expected phonetic sequence?
- Which parts of a pronunciation require further analysis?
The released implementation is therefore useful not only as a pronunciation-training system, but also as a research platform for experimenting with phonetic representations and alignment.
| Document | Description |
|---|---|
| INFO.md | Detailed explanation of the TIPAA-SSL implementation |
| API_v2.md | API endpoints documentation |
| INSTALLATION.md | Installation instructions, incl. macOS/Apple Silicon |
| MODELS.md | Pretrained model download and setup (hf_models/, quantized ONNX) |
| DATASETS.md | Speech dataset download instructions (actor recordings, LibriSpeech, MAILABS) |
| TESTING.md | How to run the tests and what they need |
| DEVELOPMENT.md | Best practices, Jupyter notebooks, TODO |
| DEPLOYMENT.md | Deployment information |
A presentation covering the two research contributions was given as part of the ISCA SIG-SLATE webinar series.
Watch the presentation on YouTube
If you use this repository in academic work, please cite the relevant paper(s).
@Article{acoustics6030042,
author = {Tits, Noé and Bhatnagar, Prernna and Dutoit, Thierry},
title = {Text-Independent Phone-to-Audio Alignment Leveraging SSL (TIPAA-SSL) Pre-Trained Model Latent Representation and Knowledge Transfer},
journal = {Acoustics},
volume = {6},
number = {3},
pages = {772--781},
year = {2024},
doi = {10.3390/acoustics6030042}
}@inproceedings{tits-2023-must,
title = {MUST\&P-SRL: Multi-lingual and Unified Syllabification in Text and Phonetic Domains for Speech Representation Learning},
author = {Tits, Noé},
booktitle = {Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track},
pages = {74--82},
year = {2023},
doi = {10.18653/v1/2023.emnlp-industry.8}
}A software citation will also be provided through the repository's CITATION.cff and archived software release.
The repository is distributed under the GNU General Public License v3.0 (GPL-3.0).
Please also consult the license and attribution information associated with individual third-party components included in or used by the project.
In particular, some dependencies and adapted components have their own licenses and attribution requirements.
This repository is a public research release of software developed for speech and language technology research.
Some parts of the original production system were developed for a larger application and may require adaptation before being used as a standalone research package.
The repository is therefore intended primarily as a research and experimentation resource, rather than as a polished production-ready speech API.
Contributions, experiments, issues, and discussion are welcome.
This work was developed in the context of research and development at Flowchase and the Numediart Institute of UMONS.
The research builds on open-source speech and language technologies including wav2vec 2.0 / XLS-R, the Montreal Forced Aligner, eSpeak, and related open-source resources.