DouK-Downloader is a tool for downloading and collecting data from Douyin and TikTok works, supporting batch downloads of videos and images, comment data collection, and live stream address acquisition with terminal, Web API, and Docker deployment options.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONspotDL is a command-line tool that finds and downloads songs from Spotify playlists via YouTube, including album art, lyrics, and metadata.
newspaper3k is a Python 3 library designed for extracting and curating articles from the web, providing tools for full-text and metadata extraction.
A GitHub Actions-based automation pipeline that daily collects recruitment notices from 42 major Korean hospitals using Selenium, verifies data integrity via a RaiT gateway, and generates filtered Excel reports.
Weibo Spider is a Python tool for crawling Sina Weibo user data. It captures user profiles and posts (text, images, videos), saving to CSV, JSON, TXT files or MySQL, MongoDB, SQLite databases, ideal for academic research and data analysis.
A lightweight API built with Elysia and FFmpeg to capture screenshots and generate GIFs from anime episode HLS links.
Scrapy is a fast, high-level web scraping framework for Python to extract structured data from websites. Cross-platform, requires Python 3.10+, maintained by Zyte and contributors.
Adaptive web scraping framework for Python: intelligent DOM tracking, anti-bot bypass fetchers, concurrent crawling with pause/resume, proxy rotation, and MCP/AI-ready output.
OCRmyPDF is a command-line tool that adds a searchable OCR text layer to scanned PDF files, producing validated PDF/A output. It uses Tesseract for recognition in over 100 languages, supports deskewing and rotation correction, and runs on Linux, macOS, Windows and FreeBSD.