OPEN SOURCE, OPEN TO EVERYONE

Open source. Open possibilities.

Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.

Human-curated · Discover open source4102discovered

A little curiosity. A world of open source.

THE FIRST COLLECTION
Topic: scraper清除
newspapercodelucas
ADDED

newspaper3k is a Python 3 library designed for extracting and curating articles from the web, providing tools for full-text and metadata extraction.

Data & databasesAutomationData integration
easyspiderNaiboWang
ADDED

EasySpider (易采集) is a free visual browser automation and web scraping tool with a no‑code GUI, command‑line mode, flowchart‑style task design, looping, conditional branches, OCR, custom JS/Python, MySQL output, scheduled and parallel runs, IP pool switching, and AGPL‑3.0 license; featured at WWW 2023 with a patent and a Zhejiang University thesis.

AutomationDeveloper toolsBrowser & RPA

A self-hosted data API for Douyin and TikTok. It fetches posts, profiles, comments, and playlists, and downloads videos and image albums without watermarks. Features a self-healing identity pool, PostgreSQL archiving, REST API, MCP server, CLI, and a web console.

Self-hostedData & databasesServer management
luxiawia002
ADDED

Lux is a fast and simple video downloader CLI tool written in Go. It supports downloading videos, playlists, images, and audio from various sites like YouTube, Bilibili, and Instagram.

Film, Video & MediaDeveloper toolsDownload & recording
huginnhuginn
ADDED

Huginn is a self-hosted system for building agents that automate online tasks by monitoring events and taking actions on your behalf.

AutomationSelf-hostedWorkflow automation
firecrawlfirecrawl
ADDED

Firecrawl is an open-source web data API for searching, scraping, crawling, and interacting with websites, turning pages into clean Markdown, JSON, or screenshots for LLM-ready use by AI agents and applications.

AI & MLDeveloper toolsAI agents
ADDED

Crawlee is a Python web scraping and browser automation library for building reliable crawlers. It supports HTTP and headless browser crawling, automatic retries, proxy rotation, and data storage for AI/LLM applications.

Developer toolsAI & MLBackend & APIs

A Bilibili video batch downloader implemented with requests, supporting multi-part downloads, subtitles, danmaku, and cover grabbing, featuring a GUI and AI corpus transcription.

Film, Video & MediaDownload & recording