VoxCPM2 is a tokenizer-free text-to-speech system with a 2B diffusion autoregressive model supporting 30 languages, voice design from text descriptions, controllable voice cloning, and 48kHz audio output.
OPEN SOURCE, OPEN TO EVERYONE
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
✳Human-curated · Discover open source3941discovered
A little curiosity. A world of open source.
THE FIRST COLLECTIONTopic: speech-synthesis清除
voxcpmOpenBMB
ADDEDspeechNVIDIA-NeMo
ADDEDNVIDIA NeMo Speech is an Apache-2.0 PyTorch framework for building, customizing and deploying speech AI models, covering automatic speech recognition, text-to-speech and speech LLMs, with pre-trained checkpoints and Docker/uv install paths.
funttsfarfarfun
ADDEDFunTTS is a text-to-speech library with a unified interface that supports seamless switching between various open-source TTS engines, offering features like subtitle generation and multi-role dialogue.
speech-to-speechhuggingface
ADDEDA low-latency, modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) that implements the OpenAI Realtime event set over WebSocket and WebRTC.