About this project
Hippo is a memory layer for AI coding agents, distributed as the npm package `hippo-memory`. Its stated premise is that good memory means knowing what to forget: memories that turned out wrong, were replaced, or were never used should rank lower or leave recall entirely. It is MIT-licensed and requires Node.js 22.16+.
Storage and retrieval
Memories live in a local SQLite database at `.hippo/hippo.db`, with markdown mirrors written after each change so the store is human-readable and git-trackable. Search is BM25 by default, with no model and no network call; embeddings are an optional install (local Transformers.js, or opt-in API embedders from OpenAI, Voyage or Cohere). The project reports zero runtime dependencies. An MCP server and an HTTP API expose the store to any compatible client.
Learning from outcomes
Two mechanisms are described as measured: marking a recalled memory as wrong lowers its ranking, and repeated use strengthens a memory. `hippo supersede` removes a replaced fact from recall. Error-tagged memories receive twice the half-life of ordinary ones. Default half-life is 365 days; the README notes a pre-registered evaluation where 7 days lost the current version of a fact far more often, while 730 days and decay-off tied with 365. Decay, three-layer storage and sleep consolidation are described as inspired by the hippocampus, and the README states that in their tests decay tied with decay disabled and sleep lowered recall.
Setup and integrations
`hippo init` run inside a project creates the `.hippo/` store, adds a marked block to an existing `CLAUDE.md` or `AGENTS.md`, installs hooks for detected agents (Claude Code settings, an OpenCode plugin, Codex `hooks.json` entries that require one-time trust), and schedules a daily 6:15am run. `hippo init --scan <folder>` finds git repos up to three levels deep and seeds each from the last 365 days of commits. Flags `--no-hooks`, `--no-schedule` and `--no-learn` skip individual parts. `hippo doctor` verifies the setup.
Features
Importers cover ChatGPT exports, CLAUDE.md, Cursor rules, markdown and plain text, with dry-run and duplicate detection. Conversation capture extracts decisions, rules, errors and preferences with pattern heuristics rather than an LLM. Slack and GitHub webhook connectors ingest events as raw memories with provenance and deletion handling. Additional commands cover active task snapshots, session event trails, handoffs, a bounded working-memory scratchpad, confidence tiers (verified, observed, inferred, stale), conflict tracking, invalidation on migration commits, and explainable recall via `hippo recall --why`.
Privacy notes
Recall makes no network call by default. The README discloses that `hippo sleep` sends memory text to Anthropic for fact extraction when `ANTHROPIC_API_KEY` is set, and that this can be disabled in `.hippo/config.json`. The hosted TypeSafe Jev reranker is opt-in and off by default.
Evidence and caveats
The README links each claim to a benchmark or test, including a sequential-learning benchmark, a LongMemEval oracle split result (R@5 = 74.0% with BM25 only), a private 300-query reranker evaluation, and a staged Slack corpus. It also states that one earlier magnitude claim was retracted, and that the supersede mechanism has not been measured. The project reports 3,500+ tests against a real database with no module mocks.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.