A Bilibili video batch downloader implemented with requests, supporting multi-part downloads, subtitles, danmaku, and cover grabbing, featuring a GUI and AI corpus transcription.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONnewspaper3k is a Python 3 library designed for extracting and curating articles from the web, providing tools for full-text and metadata extraction.
EasySpider (易采集) is a free visual browser automation and web scraping tool with a no‑code GUI, command‑line mode, flowchart‑style task design, looping, conditional branches, OCR, custom JS/Python, MySQL output, scheduled and parallel runs, IP pool switching, and AGPL‑3.0 license; featured at WWW 2023 with a patent and a Zhejiang University thesis.
A self-hosted data API for Douyin and TikTok. It fetches posts, profiles, comments, and playlists, and downloads videos and image albums without watermarks. Features a self-healing identity pool, PostgreSQL archiving, REST API, MCP server, CLI, and a web console.
Lux is a fast and simple video downloader CLI tool written in Go. It supports downloading videos, playlists, images, and audio from various sites like YouTube, Bilibili, and Instagram.
Huginn is a self-hosted system for building agents that automate online tasks by monitoring events and taking actions on your behalf.
Firecrawl is an open-source web data API for searching, scraping, crawling, and interacting with websites, turning pages into clean Markdown, JSON, or screenshots for LLM-ready use by AI agents and applications.
Crawlee is a Python web scraping and browser automation library for building reliable crawlers. It supports HTTP and headless browser crawling, automatic retries, proxy rotation, and data storage for AI/LLM applications.