Oneiron is an embedded, ACID-compliant retrieval engine for AI agents, combining vector, text, graph, temporal, and phonetic signals in a single LMDB environment.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONA lightning-fast search engine API providing AI-powered hybrid search, typo tolerance, and search-as-you-type capabilities for websites and applications.
Context7 is an MCP server and CLI that fetches up-to-date, version-specific documentation and code examples from libraries directly into LLM prompts, reducing hallucinations and outdated API suggestions for AI coding assistants.
A personal repository of articles and experiments covering statistics, machine learning, economics, and mathematics, focusing on clear, practical explanations of complex topics.
Qdrant is a high-performance vector similarity search engine and database written in Rust, designed for AI applications requiring semantic search and large-scale vector management.
Google Cloud Knowledge Catalog (formerly Dataplex) is an AI-powered data catalog and metadata management platform. This repository provides tools, agents, and samples for building context management, enrichment, and retrieval solutions using a dynamic knowledge graph of structured and unstructured data.
Faiss is a C++ library with Python wrappers for efficient similarity search and clustering of dense vectors, supporting CPU and GPU indexes, quantization, and billion-scale datasets.
A comprehensive open-source guide covering prompt engineering techniques, RAG, AI agents, and LLM optimization. Includes tutorials, research papers, tools, and notebooks for working effectively with large language models.
OpenViking is an open-source context database for AI agents that unifies memory, knowledge RAG and skills under a virtual filesystem (viking://). It offers layered context loading, directory-scoped retrieval, session-to-memory extraction, SDKs, MCP integrations and a self-hostable server.
LlamaIndex is an open-source data framework designed to build agentic applications by augmenting LLMs with private data.
Elasticsearch is a distributed, RESTful search and analytics engine and vector database. It supports full-text, vector, and hybrid search, logs, metrics, APM, security analytics, and RAG use cases. Includes quick local Docker setup, APIs, and Kibana.
LangExtract is a Python library from Google that uses LLMs to extract structured information from unstructured text, grounding every extraction to its exact source location and generating interactive HTML visualizations.
A community-maintained catalog of Bangla (Bengali) NLP resources — 712 papers, 63 datasets, 20 models, and 9 tools across 13 tasks — built as a static Astro site with strict data verification and shareable filters.
AllSource is an AI-native event store built in Rust, featuring high-throughput ingestion, low-latency reads, and native MCP integration for AI agents. It supports durable event sourcing without Postgres, includes an agent memory engine, and offers community and enterprise editions.
A framework for computing state-of-the-art text embeddings, retrieval, and reranking using Sentence Transformer, Cross-Encoder, Sparse Encoder, and Multi-Vector models, with support for training custom models.
Starbase is a web-based database and toolkit for exploring Starship transposable elements in fungi, offering sequence browsing, BLAST/HMMER search, submission, and visualization features.
A curated collection of 90+ hands-on AI engineering tutorials and projects covering LLMs, RAG, agents, MCP, multimodal apps, fine-tuning and evaluation, organized by beginner, intermediate and advanced difficulty.
Lexos is an event-driven AI document processing engine combining a Go gateway, Python workers, Redis queues and self-hosted models for offline RAG, summarization and speech transcription.
ParadeDB is a Postgres extension that integrates full-text search, vector retrieval, and analytics into a single database.
LEANN is a lightweight vector database designed for personal RAG applications, reducing storage requirements by up to 97% through graph-based selective recomputation.