vtmate is a terminal-based voice AI toolkit featuring ultra-low latency, 41 language support, and integrated TTS/STT. It enables live voice conversations, debates, voice cloning, session exports, and background daemon mode with global shortcuts for seamless AI interaction.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONOpen-source AI workspace for creators, integrating video translation, downloading, image generation, dubbing, and video generation into one local Agent environment powered by Codex CLI.
Unsloth is a desktop application and framework for running and training large language models (LLMs) and diffusion models locally across Windows, macOS, and Linux.
Verba is an offline-first desktop app for learning a language by talking with an AI coach. It runs locally with Ollama or your own provider key, and offers inline corrections, graded reading, vocabulary spaced repetition, plus bundled speech and dictation.
A browser extension for Google Chrome and Mozilla Firefox that provides real-time, dual-language subtitle translation for YouTube videos.
LocalAI is an open-source, self-hosted AI engine that runs LLMs, vision, voice, image and video models on any hardware, including CPU-only. It offers OpenAI, Anthropic and ElevenLabs-compatible APIs, on-demand modular backends, multi-user auth, and built-in agents with RAG and MCP.
VoxCPM2 is a tokenizer-free text-to-speech system with a 2B diffusion autoregressive model supporting 30 languages, voice design from text descriptions, controllable voice cloning, and 48kHz audio output.
Pixelle-Video is an AI-powered automatic short video generation engine that creates videos from a single theme, including scripts, images, voiceovers, and background music, supporting multiple AI models and workflows.
A terminal-based hands-on Python course: you write real code in a Neovim-style modal editor while a local TTS voice coaches you, offline. Includes drills, a simulated cloud/DevOps track (git, Docker, AWS, Terraform, Ansible, CI/CD) and saved progress.
ChatTTS is an open-source text-to-speech model built for dialogue scenarios, supporting English and Chinese with multi-speaker output and fine-grained control over laughter, pauses and prosody.
GPT-SoVITS is an open-source few-shot voice conversion and TTS system supporting zero-shot (5s) and few-shot (1min) voice cloning across Chinese, Japanese, Korean, Cantonese, and English. Includes WebUI tools for audio separation, ASR, and dataset preparation.
Open-source video translation, audio transcription, AI dubbing, and subtitle translation tool with local and online API support.
NVIDIA NeMo Speech is an Apache-2.0 PyTorch framework for building, customizing and deploying speech AI models, covering automatic speech recognition, text-to-speech and speech LLMs, with pre-trained checkpoints and Docker/uv install paths.
FunTTS is a text-to-speech library with a unified interface that supports seamless switching between various open-source TTS engines, offering features like subtitle generation and multi-role dialogue.