newspaper3k is a Python 3 library designed for extracting and curating articles from the web, providing tools for full-text and metadata extraction.
OPEN SOURCE, OPEN TO EVERYONE
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
✳Human-curated · Discover open source4109discovered
A little curiosity. A world of open source.
THE FIRST COLLECTIONTopic: crawling清除
newspapercodelucas
ADDEDscrapyscrapy
ADDEDScrapy is a fast, high-level web scraping framework for Python to extract structured data from websites. Cross-platform, requires Python 3.10+, maintained by Zyte and contributors.
scraplingD4Vinci
ADDEDAdaptive web scraping framework for Python: intelligent DOM tracking, anti-bot bypass fetchers, concurrent crawling with pause/resume, proxy rotation, and MCP/AI-ready output.
crawlee-pythonapify
ADDEDCrawlee is a Python web scraping and browser automation library for building reliable crawlers. It supports HTTP and headless browser crawling, automatic retries, proxy rotation, and data storage for AI/LLM applications.