About this project

OrbitKV is an external key-value cache layer for LLM inference engines. It extends the KV cache of vLLM and SGLang beyond GPU memory: reusable prefixes are kept in DRAM and SSD, then restored when a matching request arrives. The stated use cases are workloads with repeated documents, shared system prompts, and long conversations whose prefixes no longer fit in the engine's GPU cache. Deployment model An independent Cache Manager runs per host, and engines on that host connect to its shared cache. Engines keep ownership of GPU memory and scheduling; OrbitKV manages external replicas and transfers. The single-node path is described as GPU-tested on vLLM 0.29.0 and SGLang 0.5.20. Multi-node cache sharing is labeled experimental, and interfaces may change before 1.0. Key capabilities - DRAM and SSD caching: prefixes can be reused after GPU eviction or an engine restart while the Cache Manager stays alive. - Optional reuse policies: a Rust component can protect reused pages within a byte cap and admit SSD writes selectively; the documentation notes a cold-reuse tradeoff. - Direct GPU transfers: both engines register GPU buffers through CUDA IPC; adapters fence the producing CUDA stream, and Rust handles cache queries, reads and transfer completion. - Model-aware recovery: cache identity includes model artifacts, computation settings and storage layout. Compiled recovery rules select attention pages, sliding windows and recurrent/conv checkpoints, including layouts combining all three. - Bounded resource use: byte budgets cover pending reads, ready pages and active GPU transfers; cancellation retains submitted I/O until completion. - Observability: Prometheus metrics and optional request timelines, with reproduction commands for published measurements. - Experimental shared cache: embedded catalog shards locate peer replicas, Mooncake Transfer Engine moves bytes, and etcd tracks cluster membership. Quickstart outline A wheel is built and installed per Python and CUDA runtime, with separate environments for vLLM and SGLang. The wheel includes the Cache Manager and Mooncake libraries; the Manager also requires compatible PyTorch. The first Python release is described as being prepared, so package publication status should be checked in the release documentation. A Cache Manager is started with a command such as `orbitkv-cache-manager --addr 127.0.0.1:50055 --http-addr 127.0.0.1:9091 --pool-size 8gb`. vLLM is then started with prefix caching enabled and a KV transfer config naming the OrbitKV connector module; SGLang is started with an OrbitKV endpoint environment variable, a page size, and the external linker and radix cache backend options. SSD caching is enabled by adding `--ssd-cache-path` and `--ssd-cache-capacity` to the Manager command; the SSD cache is recreated when the Manager restarts. The default SSD backend tries native cuFile on supported mounts and falls back to io_uring. A `--ssd-read-path uring|cufile` override selects the demand restoration route independently of the stored representation. Storage encoding options include nvCOMP ANS lossless compression, FP8, and 3/4-bit TurboQuant, selected with `--storage-codec`; the default is exact storage, and lossy modes require model-quality qualification. A rise in the `orbitkv_load_bytes_total` metric is given as confirmation of an external restore. Architecture and status The engine adapter identifies missing state and supplies GPU destinations; OrbitKV selects compatible cached ranges, reads them from configured tiers, and retains page ownership until the GPU copy finishes. Newly computed KV is published for later reuse. The same adapter API serves DRAM, SSD and experimental remote fetches, with physical placement inside the Cache Manager. The project documents an implementation plan mapping pinned LMCache, FlexKV and Mooncake mechanisms to deployment and validation work. Bounded Rust cost observations, raw-copy shadow predictions and independent SSD read routes are described as implemented; dynamic cost selection remains planned, and observations are off by default. Cross-engine byte conversion, production catalog HA and KV-aware request routing are listed as planned work. Performance documentation Performance is stated to depend on prefix reuse, cache capacity, storage and engine scheduling. Reports cover single-node comparisons (native HBM, engine CPU caches, OrbitKV, LMCache, FlexKV), SSD recovery, ordinary recovery with Qwen3-8B host reads and GPU transfers, request preparation, and shared-cache qualification. Request preparation is off by default: the documented Qwen3-8B controls improve throughput in both engines, but SGLang P95 latency regresses, and the single-H20 measurements are not presented as a universal advantage over other caches. Benchmark code and summaries are in the `benches/` directory. Licensing and contribution OrbitKV is licensed under Apache-2.0. Documentation covers installation, adapter configuration, Manager options, metrics, fault qualification, deployment patterns, a contributor guide, Python package details, test gates and releases. Contributions are expected to include relevant checks and documentation changes.