About this project

AetherGraph is a high-performance graph sampling and feature serving library designed for GNN training on billion-scale evolving graphs. It combines a Rust core with Python bindings, enabling efficient handling of massive graphs that exceed GPU memory limits. **Key Features:** - **Three graph modes:** Static CSR graphs (memory-mapped from NVMe), dynamic C-tree graphs (single-writer ingest, lock-free readers), and heterogeneous graphs (multi-relational with typed sampling) - **High performance:** io_uring async NVMe reads, NVMe passthrough, compressed graph files (Elias-Fano + StreamVByte), half-precision features (f16/bf16), and Rabbit Order vertex reordering for cache locality - **Real-time features:** GPUDirect RDMA serving (<5us) for live feature updates, bypassing CPU - **PyTorch Geometric integration:** Drop-in replacement for NeighborLoader with standard Data/HeteroData outputs - **Distributed training:** Ray Data integration with replicated topology (graph local to each worker) - **Storage flexibility:** Out-of-core paging, shared feature memory across processes, compressed formats, and tiered caching - **Free-threaded CPython support:** No-GIL builds with Rust locks **Performance benchmarks:** 1.2-1.5x faster sampling than PyG's C++ kernel, with dynamic graph operations achieving 9M inserts/sec and 38M reads/sec. **Installation:** `pip install aethergraph` with optional PyG integration. Platform-specific features (io_uring, RDMA, GPUDirect, etc.) degrade gracefully at runtime. **License:** MIT OR Apache-2.0