About this project
GrooveSeek is an MCP server for hybrid search — semantic plus full-text — over a knowledge base of Markdown files, with plain text, PDF, Office documents and source code available as opt-ins. The command is `groove`.
Out of the box, with no config file and the default model, it indexes English Markdown with YAML frontmatter. Multilingual or Japanese content requires `--model bge-m3`; every other file format is enabled in `groove.toml`.
How it works: Markdown (and optionally .txt, .pdf, .docx, .xlsx, .pptx, plus Rust source since v1.2.0, Python since v1.3.0 and PHP since v1.5.0, whose grammars are downloaded and placed manually) is parsed with YAML frontmatter and split into heading-based chunks — or, for source code, one chunk per definition. Embeddings are generated with FastEmbed by default (BGE-small-en-v1.5, or BGE-M3 for multilingual knowledge bases) and stored in SQLite with sqlite-vec for vector similarity search. A trusted configuration can instead opt into an OpenAI-compatible embedding endpoint.
Clients connect via stdio (default, one client) or Streamable HTTP (many clients), so it works with Claude Code, Cursor or any MCP-compatible client. A live-sync file watcher keeps the index fresh on manual edits, `git pull` and external scripts. An optional TOML schema can validate frontmatter conventions through `groove validate`, and since v1.9.0 the extra keys it declares are stored by `groove index` so `groove search --field` can filter on them.
The MCP surface exposes six tools — `search`, `get_document`, `list_topics`, `get_connection_graph`, `get_best_practice`, `rebuild_index` — four prompts, and the knowledge base as `kb://` resources. With `--transport http`, the server also answers on `/ui`, an operator view showing version, document and chunk counts, embedding model, watcher state, uptime and pid, with search through the same `/mcp` endpoint clients use.
Installation is via pre-built binaries for Linux x86_64 and aarch64, macOS Apple Silicon and Windows x86_64, with SHA-256 checksums and GitHub artifact attestations; Intel Mac builds must come from source. ONNX runtime and SQLite are statically linked, and embedding models are downloaded from HuggingFace on first run. Building from source is `cargo build --release`.
Quick start is `groove index --kb-path /path/to/knowledge-base`, then an `.mcp.json` entry pointing at `groove serve --kb-path ...`, or direct queries with `groove search "..." --limit 3`. Documentation covers commands, configuration keys, client recipes, the retrieval pipeline (RRF, reranking, MMR, parent retriever), filters, citations with byte offsets, evaluation against a golden query set, architecture and stability guarantees; a Japanese version of each page is provided, and the pages are published as a site. Deployment recipes for personal stdio, NAS-shared and intranet HTTP setups are included. Releases before 1.0.0 are beta with no compatibility guarantee; from 1.0.0 the stability document states what is frozen and what is deliberately not, including the web interface and the Rust API. Dual-licensed MIT or Apache-2.0.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.