ChatMonteur is an agent-orchestrated AI video editor for talking-head footage. It transcribes, cuts by meaning, adds captions, motion graphics, sound, and color — driven by Claude Code or Codex. Two tested quality gates reject weak plans and broken files. Free and local by default.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONA Python tool for localizing authorized YouTube videos with Chinese subtitles and assisting in Bilibili uploads, supporting speech transcription, OCR, LLM translation, and hard-subtitle rendering.
Voicebox is an open-source, local-first AI voice studio that clones voices, generates speech, provides global dictation, and lets AI agents speak through MCP-aware tools.
Magnetic is a free, open-source non-linear video editor for Windows featuring a magnetic timeline and AI-assisted editing tools.
Verba is an offline-first desktop app for learning a language by talking with an AI coach. It runs locally with Ollama or your own provider key, and offers inline corrections, graded reading, vocabulary spaced repetition, plus bundled speech and dictation.
An automation runner that keeps the Windows executable published by the companion package Soenneker.Libraries.Whisper.CTranslate up to date, driven by scheduled and on-demand GitHub Actions workflows.
LearnNote is a local-first AI learning assistant that transforms videos from Bilibili, YouTube, and local files, as well as PDFs, into structured notes via subtitles or transcription, supporting reading, editing, and exporting in a web workbench.
A high-performance C/C++ implementation for inference of OpenAI's Whisper automatic speech recognition (ASR) model, designed for lightweight and cross-platform deployment.
SkyClass Distill is a pipeline tool that transforms teaching videos into structured Teaching Skills, supporting video collection, transcription, evidence extraction, and LLM distillation.
Xinference is a unified inference library for serving language, speech, and multimodal models. It provides OpenAI-compatible RESTful API, heterogeneous GPU/CPU utilization, distributed deployment, and integrates with LangChain, LlamaIndex, Dify, and Chatbox.
A browser-based subtitle editor featuring speaker-specific lanes to streamline the process from transcription to correction and export.
by2kb converts video URLs from platforms like Bilibili and YouTube into durable Markdown knowledge base artifacts, including transcripts, abstracts, and study notes, supporting agent-first workflows and local/cloud ASR.
An AI-powered automatic video editing tool using Whisper and Gemini that supports Thai and English, featuring key segment cutting and subtitling with 3-4x faster performance via GPU and Docker.
A fast reimplementation of OpenAI's Whisper model using CTranslate2, offering improved transcription speed and reduced memory usage.
Tandem is a self-hosted server with web and Android clients that keep one reading position across an ebook and its audiobook. It transcribes the audiobook with Whisper, aligns the transcript to the EPUB text sentence by sentence, and stores progress server-side.
LazyEdit is a local-first, AI-assisted video workflow for end-to-end content creation, processing, and optional publishing.
WhisperLiveKit is a self-hosted, ultra-low-latency speech-to-text pipeline. It provides real-time transcription, speaker diarization, and translation via WebSocket and OpenAI/Deepgram-compatible APIs.