SenseVoiceSmall is an open-source speech foundation model for ASR, language ID, emotion recognition and audio event detection, covering Mandarin, Cantonese, English, Japanese and Korean with fast non-autoregressive inference.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONNVIDIA NeMo Speech is an Apache-2.0 PyTorch framework for building, customizing and deploying speech AI models, covering automatic speech recognition, text-to-speech and speech LLMs, with pre-trained checkpoints and Docker/uv install paths.
Douvo is a lightweight macOS voice input app for Apple Silicon that uses Doubao ASR with optional local or remote AI post-processing for text cleanup, translation, and selected-text editing.
Batchalign3 is a TalkBank audio and ML pipeline for producing and enriching CHAT transcripts, covering ASR, forced alignment, neural morphotagging, translation and utterance segmentation. It ships a CLI, Python package, PyO3 bridge, React dashboard and an experimental Tauri desktop shell.
A Python library that retrieves YouTube video transcripts and subtitles without requiring an API key or headless browser. Supports multiple languages, translation, and output formatting.
Mimora is a free, fully local pronunciation trainer for English and Spanish. It uses on-device TTS, speech recognition, and a local LLM to generate phrases and score user attempts with word-level feedback, requiring no cloud or API keys.
Premove ITN is an open-source, context-aware inverse text normalization tool for English voice-agent transcripts. It converts spoken ASR output into canonical written forms using a generate-score-decode pipeline with DeBERTa contextual scoring.