ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
-
Updated
Sep 17, 2026 - Python
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
AI Speech Solutions for Tasks such as ASR, Vocal Extraction, Accompaniment Extraction, Audio Denoising, and Enhancement, Support models such as paraformer, sensevoice, fireredasr, zipformer, moonshine, wenet, whisper, fsmn-vad, silero-vad, CT Transformer punc, Spleeter, Uvr5, etc, apply ONNX models in various scenarios.
High-throughput batched ASR for NVIDIA GPUs. Run Zipformer and NVIDIA NeMo Parakeet models through one Python API with TensorRT, custom CUDA plugins, GPU beam search, and word timestamps - up to 25,000 RTFx on B300.
An upgrade framework for train and validate compare with icefall using Lightning.
Fast multilingual speech-to-text HTTP server. 8 languages, Silero VAD, word timestamps. CPU-only, Rust + sherpa-onnx.
Rust project. Local Free AI transcription based on whisper, parakeet, MoonShine, SenseVoice, Zipformer.
Fast Speech-to-Text architectures with bindings to all languages
Streaming speech-to-text library for Rust with a unified trait interface over multiple backends (sherpa-onnx Zipformer/Paraformer). Part of the WaveKat voice pipeline.
Local Vietnamese STT server powered by Zipformer-vi RNN-T
A template for serving zipformer on Triton Inference Server.
Offline Speech-to-Text for React Native using sherpa-onnx Supports Zipformer, Paraformer, NeMo CTC, Whisper & more.
A lightweight smart meeting assistant. Integrating speech recognition and AI summarization to streamline the note-taking process and automatically generate efficient meeting minutes and action items.
Local CPU WebSocket and browser demo for icefall streaming ASR ONNX models
Generic on-device speech-to-text. Local inference, no cloud APIs, full privacy. Rust + ONNX Runtime.
Streaming piano-transcription system
Local-first Vietnamese/English streaming speech-to-text PWA with private voice personalization
🎤 Enable offline speech recognition in React Native using sherpa-onnx, supporting various model architectures for reliable performance.
To associate your repository with the zipformer topic, visit your repo's landing page and select "manage topics."