SenseVoiceSmall is an open-source speech foundation model for ASR, language ID, emotion recognition and audio event detection, covering Mandarin, Cantonese, English, Japanese and Korean with fast non-autoregressive inference.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONA local multimodal LLM plugin for Unreal Engine 5 providing offline inference for text, speech-to-text, and text-to-speech.
Jano is a lightweight OpenAI-compatible router that enables multiple local LLMs to share one GPU by intelligently batching requests and minimizing model swaps.
Lexos is an event-driven AI document processing engine combining a Go gateway, Python workers, Redis queues and self-hosted models for offline RAG, summarization and speech transcription.
ELI is a strictly local, private AI assistant that runs entirely on your own hardware. It offers voice, vision, memory, and 227 capabilities across 17 areas, including computer control, media playback, document handling, coding, and task automation. Offline by default, it uses any local GGUF model and is model-agnostic, with optional GPU acceleration and a web dashboard for remote LAN access.
vtmate is a terminal-based voice AI toolkit featuring ultra-low latency, 41 language support, and integrated TTS/STT. It enables live voice conversations, debates, voice cloning, session exports, and background daemon mode with global shortcuts for seamless AI interaction.
MoziAI-35B is a locally deployable open-source multimodal LLM with a 35B MoE architecture compressed to ~15.9 GB via MoziSmartBit quantization, offering 256K context, vision, tool calling, and a finance-focused reasoning framework.
Mimora is a free, fully local pronunciation trainer for English and Spanish. It uses on-device TTS, speech recognition, and a local LLM to generate phrases and score user attempts with word-level feedback, requiring no cloud or API keys.
Xinference is a unified inference library for serving language, speech, and multimodal models. It provides OpenAI-compatible RESTful API, heterogeneous GPU/CPU utilization, distributed deployment, and integrates with LangChain, LlamaIndex, Dify, and Chatbox.