SoundsGoodAI builds open-source infrastructure for fast, accurate, and efficient speech processing.
Batched offline speech recognition for NVIDIA GPUs. Raw audio in, text and word timestamps out through one Python API.
It supports Zipformer and NVIDIA Parakeet models using TensorRT engines, native CUDA plugins, GPU feature extraction, and batched GPU decoding.
High-performance Silero voice activity detection for CPUs. It combines a streamlined ONNX graph with an optimized spectral frontend to improve throughput while preserving numerically equivalent speech probabilities.
It supports Linux and macOS on both x86-64 and ARM64 systems.