Voicebox is an open-source, local-first AI voice studio that clones voices, generates speech, provides global dictation, and lets AI agents speak through MCP-aware tools.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONPremove ITN is an open-source, context-aware inverse text normalization tool for English voice-agent transcripts. It converts spoken ASR output into canonical written forms using a generate-score-decode pipeline with DeBERTa contextual scoring.
RunAnywhere is a cross-platform SDK suite for running AI models fully on-device across phones, browsers, desktops, and servers. It supports LLMs, vision, speech, voice agents, RAG, embeddings, and image generation with a single semantic API.
SenseVoiceSmall is an open-source speech foundation model for ASR, language ID, emotion recognition and audio event detection, covering Mandarin, Cantonese, English, Japanese and Korean with fast non-autoregressive inference.