AI-powered template that clones any website into a clean Next.js app via a single command. Supports Claude Code, Cursor, Codex, Gemini, and more.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONweb-core is a Python package providing shared web infrastructure: SearXNG search with retry, deduplication and domain filtering; a multi-strategy scraping agent with automatic escalation; an SSRF-safe HTTP client with DNS pinning; stealth and remote browser rendering; and typed Google Drive and MangaDex adapters.
The @knownagents/sdk is a Node.js library that lets server‑side applications track AI agent traffic, record LLM referrals, generate up‑to‑date robots.txt files, and identify bots via the Agent Identification API. It supports batch event flushing, MCP call tracking, and integration with ACP/UCP commerce flows.
A self-hosted data API for Douyin and TikTok. It fetches posts, profiles, comments, and playlists, and downloads videos and image albums without watermarks. Features a self-healing identity pool, PostgreSQL archiving, REST API, MCP server, CLI, and a web console.
Converts docs, GitHub repos, PDFs, videos into AI skills for Claude, Gemini, OpenAI with conflict detection and 22 export targets.
designlang is a Playwright-based CLI that reads a website's live DOM and extracts its design system: DTCG tokens, Tailwind config, Figma variables, shadcn theme, component anatomy, motion tokens, brand voice, WCAG contrast audits, and multi-platform emitters, plus MCP server and Chrome extension.
Scrapy is a fast, high-level web scraping framework for Python to extract structured data from websites. Cross-platform, requires Python 3.10+, maintained by Zyte and contributors.
Adaptive web scraping framework for Python: intelligent DOM tracking, anti-bot bypass fetchers, concurrent crawling with pause/resume, proxy rotation, and MCP/AI-ready output.
Firecrawl is an open-source web data API for searching, scraping, crawling, and interacting with websites, turning pages into clean Markdown, JSON, or screenshots for LLM-ready use by AI agents and applications.
Crawlee is a Python web scraping and browser automation library for building reliable crawlers. It supports HTTP and headless browser crawling, automatic retries, proxy rotation, and data storage for AI/LLM applications.
Ferret is a declarative data automation language and Go runtime for structured extraction workflows, supporting embedding, capability-based access, and cross-source querying.
Automated job alert pipeline that polls 6 major company career APIs and ~1,200 ATS job boards, filters for US-based entry/mid-level software engineering roles, deduplicates against persisted state, and emails digest alerts on a schedule via GitHub Actions.
A personal automation that books gym classes the moment their 72-hour booking window opens, using a fast direct API call with a Playwright browser-automation fallback, plus a schedule scraper feeding a GitHub Pages dashboard.
An open-source tool for monitoring website changes and receiving real-time alerts via various notification channels.
J.A.R.V.I.S is a local-first OSINT platform that uses AI and web scraping to build comprehensive intelligence dossiers on individuals while ensuring data privacy.
A Python tool that scrapes the i2iFunding P2P lending marketplace every few minutes, scores loans by yield and credit, sends Telegram/ntfy alerts, and serves a static GitHub Pages dashboard backed by committed JSON files. An optional, safe-gated auto-investor can place real money.