Video2X is a machine learning-based video super resolution and frame interpolation framework, rewritten in C/C++ for cross-platform Windows and Linux support.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONQwen3-VL is Alibaba Qwen team's multimodal vision-language model series, offering Dense and MoE variants with Instruct and Thinking editions, 256K context, OCR, visual agents, and video understanding.
LivePortrait is an open-source PyTorch implementation for efficient portrait animation, supporting humans, cats and dogs. It offers image and video driven modes, stitching and retargeting control, motion templates, a Gradio interface, and pretrained weights.
ChatTTS is an open-source text-to-speech model built for dialogue scenarios, supporting English and Chinese with multi-speaker output and fine-grained control over laughter, pauses and prosody.
Tandem is a self-hosted server with web and Android clients that keep one reading position across an ebook and its audiobook. It transcribes the audiobook with Whisper, aligns the transcript to the EPUB text sentence by sentence, and stores progress server-side.
olmOCR is a toolkit for converting PDFs and images into clean, readable markdown or plain text using vision-language models. It supports local GPU inference, external vLLM servers, and cloud providers, with features like natural reading order, filtering, and parallel processing. Includes a benchmark suite and Docker support for large-scale document linearization in LLM datasets.
LazyEdit is a local-first, AI-assisted video workflow for end-to-end content creation, processing, and optional publishing.
Bob is a macOS menu bar translation and OCR tool, supporting selection, screenshot, and input translation, as well as screenshot OCR, silent recognition, QR code recognition, with integration of multiple translation and OCR services, offline recognition, and speech synthesis.
Open-source video translation, audio transcription, AI dubbing, and subtitle translation tool with local and online API support.
HunyuanVideo is an open-source large-scale text-to-video generation model framework by Tencent, featuring a unified image-video architecture, 3D VAE, and prompt rewriting capabilities.
Wan2.2 is an open-source large-scale video generation model suite supporting text-to-video, image-to-video, speech-to-video, and character animation tasks with MoE architecture and 720P output.
Handy is a free, open-source, cross-platform speech-to-text desktop app that runs entirely offline. Press a shortcut, speak, and get transcribed text pasted into any text field using Whisper or Parakeet models.
Lamitype is a free, open-source, privacy-first macOS speech-to-text app with local Qwen3-ASR on Apple Silicon, Traditional Chinese-first output, live captions, translation, proofreading, and optional cloud dictation.
Wan2.1 is an open-source suite of video foundation models supporting text-to-video, image-to-video, video editing, and multilingual visual text generation.
An AI/ML project that uses computer vision to detect animals such as elephants, cattle, and monkeys in agricultural areas, aiming to help prevent crop damage. The README is brief and directs users to the codebase for setup details.
TextBlob is a Python library for natural language processing, offering a simple API for sentiment analysis, part-of-speech tagging, noun phrase extraction, classification, tokenization, and more, built on NLTK and pattern.
Marker is a Python tool that converts PDFs, images, PPTX, DOCX, XLSX, HTML and EPUB files into Markdown, JSON, HTML or chunked output. It formats tables, equations, forms and code blocks, extracts images, and can optionally use LLMs to boost accuracy.
RuView turns WiFi Channel State Information from cheap ESP32 sensors into camera-free sensing: presence and occupancy, breathing and heart rate, falls, activity, and 17-keypoint pose. It runs entirely on edge hardware and plugs into Home Assistant, Apple Home, Google Home and Alexa via MQTT or Matter.
Upscayl is a free and open-source AI image upscaler for Linux, macOS, and Windows that enlarges and enhances low-resolution images without losing quality.
Browser-based face detection demo built on face-api.js and TensorFlow.js, packaged as an nginx-served Docker image, alongside research notes comparing Viola-Jones, HOG+SVM and CNN detectors.