About this project
Datahike is a durable Datalog database offering Datomic-compatible APIs and git-like semantics. Built on persistent data structures with structural sharing, database snapshots are immutable values that can be held, shared, and queried without locks or copying.
Key capabilities described in the README:
- Distributed Index Space: readers access persistent indices directly for read scaling without database connections.
- Flexible storage via konserve: File, LMDB, S3, JDBC, Redis, and IndexedDB backends.
- Cross-platform support: JVM, Node.js, and browser, with Clojure, ClojureScript, JavaScript, and Java APIs.
- Real-time sync over WebSocket with Kabel for browser-to-server communication.
- Time-travel: query any historical state with a full transaction audit trail.
- Data purging for regulatory compliance (GDPR-ready).
- Reported production use with billions of datoms, including deployment in government services.
The project argues for Datalog over SQL for relationship-heavy data: pattern matching over relationships, transitive and recursive queries, and multi-database joins are expressed naturally. Immutability treats the database as an append-only log of facts, enabling audit trails, debugging through time-travel, and data excision.
Usage centers on a Clojure API (datahike.api) with create-database, connect, transact, q, history, release, and delete-database. The API aims to be a drop-in replacement for a subset of Datomic on the JVM. Configuration covers storage backends and schema flexibility.
Beta language bindings include JavaScript (npm package), a GraalVM native CLI tool (dthk), a Babashka pod, a Java API with a fluent builder, C/C++ native bindings (libdatahike), and Python bindings. ClojureScript support covers Node.js (file backend) and browsers (IndexedDB with TieredStore).
Example projects include Beleg, an invoice and CRM system with web UI and LaTeX PDF generation. Production users cited include the Swedish Public Employment Service (JobTech Taxonomy, 40,000+ concepts, migrated from Datomic), Stub (accounting for small businesses in South Africa), and Heidelberg University for internal emotion tracking research.
Related projects: Proximum provides an HNSW vector index for semantic search and RAG on Datahike's persistent model; pg-datahike embeds a PostgreSQL-compatible adapter (wire protocol, SQL translator, pg_* catalogs) so PostgreSQL clients can connect without a PostgreSQL install, with bidirectional datom-layer integration and pg_dump round-tripping (beta).
Datahike is part of the replikativ ecosystem and is composed of reusable libraries: konserve (pluggable key-value storage), kabel (WebSocket transport), hasch (content-addressable hashing), incognito (extensible serialization), superv.async (core.async supervision), replikativ (CRDT synchronization), and distributed-scope (remote invocation).
Roadmap decisions are community-driven through GitHub Discussions idea upvotes. Commercial support is offered. Licensed under the Eclipse Public License.
Comments
0 people shared their preference · Deer Point appears after 10 participants
Sign in to join the discussion.