dots.ocr is a multilingual document layout parsing vision-language model that extracts text, tables, formulas and layout from images and PDFs, and can convert charts and diagrams into SVG code.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONTesseract is an open-source OCR engine that provides a command-line tool and a library for recognizing text in images across more than 100 languages.
Fooocus is an offline, open-source image generation tool built on Stable Diffusion XL, designed to produce high-quality images from simple prompts with minimal setup and no manual parameter tuning.
MoneyPrinterTurbo is an all-in-one open-source AI short video generator built on Python 3.11+. Input a theme or keyword to automatically produce HD videos in 9:16, 16:9, and 1:1 formats with script generation, voiceover, footage matching, subtitles, and editing. Supports WebUI, API, CLI, and AI Agent usage modes.
CogVideoX is an open-source video generation model series supporting text-to-video, image-to-video, and video continuation tasks. It includes 2B and 5B parameter models with diffusers and SAT inference frameworks.
YuE2 is an open music generation model that unifies symbolic planning with audio synthesis: it turns lyrics and a style prompt into an editable melody-and-chord score, then renders a full song with vocals and accompaniment. It also supports zero-shot covers and agentic score editing.
Video-subtitle-remover (VSR) is an AI-based tool for removing hardcoded video subtitles at original resolution, with inpainting repair, subtitle extraction, and batch image watermark removal. It offers both GUI and CLI interfaces.
PageLedger is a Python CLI library for auditable, page-level OCR and document extraction. It records provenance, quality signals, and budgets for each extracted page, supports resumable runs, review queues, and reruns with human-in-the-loop verification.
A comprehensive web interface for Stable Diffusion implemented using the Gradio library, providing a wide array of tools for AI image generation and manipulation.
Meta FAIR 的 SAM 2 基础模型官方代码库,支持图像与视频的可提示分割,提供模型检查点、推理 API、训练与微调代码、Web 演示及 SA-V 数据集说明。
Rembg is a versatile tool for removing image backgrounds, available as a CLI, Python library, HTTP server, or Docker container.
UMADock is an open-source molecular docking tool that places small molecules into protein binding sites using machine-learning interatomic potentials, with automated site preparation and cloud execution.
Voicebox is an open-source, local-first AI voice studio that clones voices, generates speech, provides global dictation, and lets AI agents speak through MCP-aware tools.
SpeechRecognition is a Python library for speech recognition supporting multiple engines and APIs, both online and offline, including CMU Sphinx, Google, Azure, IBM, Vosk, Whisper, and OpenAI.
VideoLingo is an all-in-one AI-powered video translation, localization, and dubbing tool that generates Netflix-quality single-line subtitles with seamless voice dubbing across multiple languages.
Tesseract.js is a JavaScript OCR library that extracts text from images in over 100 languages. It runs in browsers via WebAssembly and on Node.js, wrapping the Tesseract engine without modifying its recognition model.
A VST3/AU plugin that integrates ElevenLabs sound generation directly into your DAW for creating and slicing sound effects.
FramePack is a next-frame video diffusion model that compresses input contexts to constant length, enabling 13B model video generation on laptop GPUs with as little as 6GB VRAM.
Waifu2x-Extension-GUI is a Windows application for AI-based image, GIF, and video super-resolution upscaling and video frame interpolation, supporting AMD, NVIDIA, and Intel GPUs.
EasyOCR is a ready-to-use optical character recognition library supporting 80+ languages and major writing scripts. It offers a simple Python API and CLI, with PyTorch-based detection and recognition models, plus options for custom training.