NVIDIA NeMo Speech is an Apache-2.0 PyTorch framework for building, customizing and deploying speech AI models, covering automatic speech recognition, text-to-speech and speech LLMs, with pre-trained checkpoints and Docker/uv install paths.
AI & MLMusicModel runtime
ESPnet is an end-to-end speech processing toolkit built on PyTorch, covering ASR, TTS, speech translation, enhancement, diarization, and more, with reproducible recipes and pretrained models.
AI & MLDeveloper toolsModel runtime
A low-latency, modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) that implements the OpenAI Realtime event set over WebSocket and WebRTC.
AI & MLDeveloper toolsAI assistants