A local multimodal LLM plugin for Unreal Engine 5 providing offline inference for text, speech-to-text, and text-to-speech.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONFunTTS is a text-to-speech library with a unified interface that supports seamless switching between various open-source TTS engines, offering features like subtitle generation and multi-role dialogue.
MoneyPrinterTurbo is an all-in-one open-source AI short video generator built on Python 3.11+. Input a theme or keyword to automatically produce HD videos in 9:16, 16:9, and 1:1 formats with script generation, voiceover, footage matching, subtitles, and editing. Supports WebUI, API, CLI, and AI Agent usage modes.
Mimora is a free, fully local pronunciation trainer for English and Spanish. It uses on-device TTS, speech recognition, and a local LLM to generate phrases and score user attempts with word-level feedback, requiring no cloud or API keys.
Unsloth is a desktop application and framework for running and training large language models (LLMs) and diffusion models locally across Windows, macOS, and Linux.
Zara is an experimental local-first AI assistant and Linux desktop copilot that combines symbolic logic (Prolog) with LLMs for intent handling and reasoning.
VoxCPM2 is a tokenizer-free text-to-speech system with a 2B diffusion autoregressive model supporting 30 languages, voice design from text descriptions, controllable voice cloning, and 48kHz audio output.
An MCP server designed to teach AI musicality through piano and guitar. It features 120 annotated songs, six sound engines, a browser-based composition cockpit, and a public dataset of tool-use traces.
Velo AI Studio is an open-source, browser-based video creation tool designed to assemble narrated videos from scripts, audio, and AI-generated visual assets.
TaleWeaver is a self-hosted browser-based AI text-adventure engine with an AI gamemaster, adventure editor, RPG mechanics, voice narration/speech input, and Docker/SQLite deployment.
OpenMontage is an open-source, agentic video production system that transforms AI coding assistants into full video studios, supporting end-to-end pipelines from research to final render.
ChatTTS is an open-source text-to-speech model built for dialogue scenarios, supporting English and Chinese with multi-speaker output and fine-grained control over laughter, pauses and prosody.
GPT-SoVITS is an open-source few-shot voice conversion and TTS system supporting zero-shot (5s) and few-shot (1min) voice cloning across Chinese, Japanese, Korean, Cantonese, and English. Includes WebUI tools for audio separation, ASR, and dataset preparation.
Open-source video translation, audio transcription, AI dubbing, and subtitle translation tool with local and online API support.
ESPnet is an end-to-end speech processing toolkit built on PyTorch, covering ASR, TTS, speech translation, enhancement, diarization, and more, with reproducible recipes and pretrained models.