OPEN SOURCE, OPEN TO EVERYONE

Open source. Open possibilities.

Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.

Human-curated · Discover open source4101discovered

A little curiosity. A world of open source.

THE FIRST COLLECTION
Topic: python清除
ADDED

DouK-Downloader is a tool for downloading and collecting data from Douyin and TikTok works, supporting batch downloads of videos and images, comment data collection, and live stream address acquisition with terminal, Web API, and Docker deployment options.

AutomationSelf-hostedData automation
ADDED

spotDL is a command-line tool that finds and downloads songs from Spotify playlists via YouTube, including album art, lyrics, and metadata.

MusicAutomationTagging & conversion
newspapercodelucas
ADDED

newspaper3k is a Python 3 library designed for extracting and curating articles from the web, providing tools for full-text and metadata extraction.

Data & databasesAutomationData integration

A GitHub Actions-based automation pipeline that daily collects recruitment notices from 42 major Korean hospitals using Selenium, verifies data integrity via a RaiT gateway, and generates filtered Excel reports.

AutomationDeveloper toolsData automation
ADDED

Weibo Spider is a Python tool for crawling Sina Weibo user data. It captures user profiles and posts (text, images, videos), saving to CSV, JSON, TXT files or MySQL, MongoDB, SQLite databases, ideal for academic research and data analysis.

AutomationDeveloper toolsData automation
episodesrocky18313
ADDED

A lightweight API built with Elysia and FFmpeg to capture screenshots and generate GIFs from anime episode HLS links.

Film, Video & MediaAutomationStreaming platforms
scrapyscrapy
ADDED

Scrapy is a fast, high-level web scraping framework for Python to extract structured data from websites. Cross-platform, requires Python 3.10+, maintained by Zyte and contributors.

AutomationData & databasesData automation
scraplingD4Vinci
ADDED

Adaptive web scraping framework for Python: intelligent DOM tracking, anti-bot bypass fetchers, concurrent crawling with pause/resume, proxy rotation, and MCP/AI-ready output.

AutomationData & databasesData automation
ocrmypdfocrmypdf
ADDED

OCRmyPDF is a command-line tool that adds a searchable OCR text layer to scanned PDF files, producing validated PDF/A output. It uses Tesseract for recognition in over 100 languages, supports deskewing and rotation correction, and runs on Linux, macOS, Windows and FreeBSD.

ProductivityAutomationUtilities