About this project

Semble is a code search library built for coding agents. It lets agents query a codebase in natural language and receive only the relevant code snippets, avoiding grep-and-read workflows that consume large context windows. It can be used as an MCP server, a standalone CLI, or a Python library, and integrates with agents such as Claude Code, Cursor, Codex, and OpenCode. Key capabilities include fast CPU-only indexing and querying, local and remote repository support, multi-repo search, code/docs/config content filtering, and related-code discovery. Indexes are cached and incrementally updated when files change. File selection respects .gitignore and .sembleignore, with options to force-include non-default extensions. The retrieval pipeline combines static Model2Vec embeddings with BM25 lexical matching, fused via Reciprocal Rank Fusion and reranked with code-aware signals such as definition boosts, identifier stemming, file coherence, and noise penalties. The project reports benchmark results showing retrieval quality comparable to a code-specialized transformer while indexing and querying substantially faster, and token savings of about 99% compared with grep+read at equivalent recall. Semble is distributed as a Python package under the MIT license. It requires no API keys, GPU, or external services, though the embedding model is downloaded from Hugging Face on first use. A custom Model2Vec-compatible model can be configured through an environment variable.