About this project
Crawlee is a web scraping and browser automation library for Node.js, written in TypeScript and distributed as the `crawlee` NPM package. It aims to cover crawling and scraping end to end, providing a single interface for both HTTP crawling and real browser crawling.
Key capabilities described in the README:
- Single interface for HTTP and headless browser crawling
- Persistent queue for URLs to crawl, supporting breadth-first and depth-first ordering
- Pluggable storage for tabular data and files, defaulting to a local `./storage` directory
- Automatic scaling with available system resources
- Integrated proxy rotation and session management
- Lifecycles customizable with hooks
- A CLI to bootstrap projects (`npx crawlee create my-crawler`)
- Configurable routing, error handling, and retries
- Dockerfiles ready to deploy
- Written in TypeScript with generics
HTTP crawling features include zero-config HTTP2 support (even for proxies), automatic generation of browser-like headers, replication of browser TLS fingerprints, integrated fast HTML parsers (Cheerio and JSDOM), and the ability to scrape JSON APIs.
Real browser crawling features include JavaScript rendering and screenshots, headless and headful support, zero-config generation of human-like fingerprints, automatic browser management, and the ability to use Playwright and Puppeteer through the same interface across Chrome, Firefox, Webkit, and others.
Installation requires Node.js 16 or higher. The recommended quick start uses the Crawlee CLI, while manual installation involves adding `crawlee` and a browser automation library such as `playwright` to an existing project. A short example shows a `PlaywrightCrawler` that logs page titles, pushes results to a dataset, and enqueues links. Beta builds are published as `crawlee@next`.
The project is developed by Apify, is open-source, and runs anywhere, with documented support for deployment on the Apify platform. It is licensed under the Apache License 2.0. Support channels include GitHub issues, Stack Overflow, GitHub Discussions, and a Discord server. A Python counterpart, Crawlee for Python, is also referenced.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.