VoxCPM2 is a tokenizer-free text-to-speech system with a 2B diffusion autoregressive model supporting 30 languages, voice design from text descriptions, controllable voice cloning, and 48kHz audio output.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONNVIDIA NeMo Speech is an Apache-2.0 PyTorch framework for building, customizing and deploying speech AI models, covering automatic speech recognition, text-to-speech and speech LLMs, with pre-trained checkpoints and Docker/uv install paths.
ESPnet is an end-to-end speech processing toolkit built on PyTorch, covering ASR, TTS, speech translation, enhancement, diarization, and more, with reproducible recipes and pretrained models.
FunTTS is a text-to-speech library with a unified interface that supports seamless switching between various open-source TTS engines, offering features like subtitle generation and multi-role dialogue.
A low-latency, modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) that implements the OpenAI Realtime event set over WebSocket and WebRTC.