PixelRAG renders documents as screenshots and retrieves over images directly, preserving visual structure like tables and charts that text parsing loses. Includes a hosted 8.28M-page Wikipedia index, CLI tools, and Claude Code integration.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONTARS is a multimodal AI agent stack from ByteDance with two projects: Agent TARS, a CLI and Web UI agent integrating MCP tools and browser control, and UI-TARS Desktop, a native GUI agent app that controls computers via vision-language models.
PageLedger is a Python CLI library for auditable, page-level OCR and document extraction. It records provenance, quality signals, and budgets for each extracted page, supports resumable runs, review queues, and reruns with human-in-the-loop verification.
A growing collection of 61 computer vision tutorials covering state-of-the-art models like YOLO11, SAM 3, RF-DETR, and Qwen3-VL for object detection, segmentation, OCR, and more.
RunAnywhere is a cross-platform SDK suite for running AI models fully on-device across phones, browsers, desktops, and servers. It supports LLMs, vision, speech, voice agents, RAG, embeddings, and image generation with a single semantic API.
Hugging Face Transformers is a model-definition framework for machine learning across text, vision, audio, and multimodal tasks. It provides a unified API for inference and training with over 1 million pretrained checkpoints.
SGLang is a high-performance serving framework for large language models (LLMs) and multimodal models, designed for low-latency and high-throughput inference.