About this project

wigolo is a local-first web intelligence tool that gives AI agents a unified surface for web-related tasks: search, fetch, crawl, extract, cache, find-similar, research, and autonomous gather loops. It runs as an MCP server alongside coding agents, as a REST endpoint for self-hosted agents, or embedded via SDKs. Core tools require no API keys, all data stays under ~/.wigolo/, and there is no per-query cost. Supported agents include Claude Code, Cursor, Codex, Gemini CLI, OpenCode, VS Code, Windsurf, Zed, and Antigravity. Any MCP client can register it via npx. Framework integrations exist for LangChain, CrewAI, LlamaIndex, and Vercel AI SDK. Key tools: - search: Multi-engine web search (18 direct adapters) with rank fusion, ML reranking, and explainable per-result scoring. Supports query arrays for parallel breadth, domain/time scoping, and image results. - fetch: Tiered router escalating from plain HTTP to headless browser on anti-bot challenges or SPA shells. Returns clean markdown, metadata, links. Handles PDFs, authenticated sessions, and page actions (click/type/scroll/screenshot). - crawl: Multi-page crawl (BFS, DFS, sitemap, or map-only) with per-domain rate limits, robots.txt respect, and boilerplate dedup. - extract: Structured data extraction — tables, metadata, JSON-LD, named schemas (Article/Recipe/Product), or custom JSON Schema. - cache: Query previously seen results via keyword or hybrid semantic search, with stats, clear, and change detection. - find_similar: Pages similar to a URL or concept via 3-way fusion of keyword, semantic, and live web signals. - research: Decomposes a question into sub-queries, fetches sources, and synthesizes a cited report. - agent: Autonomous gather loop with plan/search/fetch/extract/synthesize, step log, time budget, and optional output schema. - diff + watch: Track page changes since last visit, re-check on demand, deliver to webhook. Every result carries a verbatim excerpt pinned to byte-offset source spans, a citation ID, and an evidence score. Weak results are flagged as junk; failed engines and stale cache are labeled. The fetch router learns per-domain when to escalate to a real browser and unlearns when a site stops needing it. Setup via npx wigolo init (with optional --agents flag to wire specific agents). Requires Node >= 20 and ~1.5 GB disk on macOS, Linux, or Windows. Docker images available (slim and :full with preinstalled browser engine). TypeScript and Python SDKs provided. An 11-pack skill catalog teaches coding agents to use each tool. Optional LLM provider (Gemini, Anthropic, OpenAI, Groq, or local Ollama) enables synthesized cited answers for research/agent tools; without one, raw briefs and evidence are returned for the host LLM to assemble. Licensed under AGPL-3.0, currently in public beta.