Douvo is a lightweight macOS voice input app for Apple Silicon that uses Doubao ASR with optional local or remote AI post-processing for text cleanup, translation, and selected-text editing.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONLexos is an event-driven AI document processing engine combining a Go gateway, Python workers, Redis queues and self-hosted models for offline RAG, summarization and speech transcription.
SpeechRecognition is a Python library for speech recognition supporting multiple engines and APIs, both online and offline, including CMU Sphinx, Google, Azure, IBM, Vosk, Whisper, and OpenAI.
A high-performance C/C++ implementation for inference of OpenAI's Whisper automatic speech recognition (ASR) model, designed for lightweight and cross-platform deployment.
Bolo is a free, self-hosted macOS push-to-talk dictation app written in Rust: hold a hotkey, speak, and the transcription is pasted at your cursor via Telnyx AI speech-to-text, with optional LLM cleanup and local transcript history.
SkyClass Distill is a pipeline tool that transforms teaching videos into structured Teaching Skills, supporting video collection, transcription, evidence extraction, and LLM distillation.
TypeNo is a free, open-source macOS voice input app that transcribes speech locally and pastes text instantly — no accounts, no settings, privacy-first.
Velo AI Studio is an open-source, browser-based video creation tool designed to assemble narrated videos from scripts, audio, and AI-generated visual assets.
A fast reimplementation of OpenAI's Whisper model using CTranslate2, offering improved transcription speed and reduced memory usage.
SenseVoiceSmall is an open-source speech foundation model for ASR, language ID, emotion recognition and audio event detection, covering Mandarin, Cantonese, English, Japanese and Korean with fast non-autoregressive inference.
LazyEdit is a local-first, AI-assisted video workflow for end-to-end content creation, processing, and optional publishing.
Open-source video translation, audio transcription, AI dubbing, and subtitle translation tool with local and online API support.
Handy is a free, open-source, cross-platform speech-to-text desktop app that runs entirely offline. Press a shortcut, speak, and get transcribed text pasted into any text field using Whisper or Parakeet models.
WhisperLiveKit is a self-hosted, ultra-low-latency speech-to-text pipeline. It provides real-time transcription, speaker diarization, and translation via WebSocket and OpenAI/Deepgram-compatible APIs.
Lamitype is a free, open-source, privacy-first macOS speech-to-text app with local Qwen3-ASR on Apple Silicon, Traditional Chinese-first output, live captions, translation, proofreading, and optional cloud dictation.
A local multimodal LLM plugin for Unreal Engine 5 providing offline inference for text, speech-to-text, and text-to-speech.
A low-latency, modular voice-agent pipeline (VAD -> STT -> LLM -> TTS) that implements the OpenAI Realtime event set over WebSocket and WebRTC.