Dagster is a cloud-native data pipeline orchestrator designed for the development, production, and observation of data assets throughout their entire lifecycle.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONPrefect is a Python-based workflow orchestration framework designed to build, automate, and monitor resilient data pipelines with features like scheduling and retries.
Apache ECharts is a free, powerful JavaScript library for creating interactive, highly customizable charts and data visualizations in the browser.
InfluxDB 3 Core is an open-source time series database designed for real-time events, analytics, and monitoring, powered by Apache Arrow, DataFusion, and Parquet.
TimescaleDB is a PostgreSQL extension designed for high-performance real-time analytics on time-series and event data.
Valkey is a high-performance, distributed key-value data structure server optimized for caching and real-time workloads, forked from the open-source Redis project.
Apache Cassandra is a highly-scalable, transactional distributed partitioned row store designed for fault-tolerance on commodity hardware or cloud infrastructure.
Polars is an extremely fast analytical query engine for DataFrames written in Rust, featuring multi-threaded execution and support for datasets larger than RAM.
Apache DataFusion is an extensible query engine written in Rust that uses Apache Arrow as its in-memory format, providing SQL and DataFrame APIs for building fast analytic systems.
Apache Beam is a unified programming model for defining both batch and streaming data-parallel processing pipelines.
Apache Doris is an open-source, real-time analytics and search database based on MPP architecture, providing fast SQL analytics and hybrid search capabilities.
OpenSearch is an open-source, enterprise-grade distributed search and observability suite designed to manage unstructured data at scale.
Milvus is a high-performance, cloud-native vector database designed for scalable approximate nearest neighbor (ANN) search, powering AI applications with unstructured data.
Apache Arrow is a universal columnar memory format and multi-language toolbox designed for efficient data interchange and in-memory analytics.
Apache Iceberg is a high-performance open table format for huge analytic datasets, bringing SQL-like reliability and simplicity to big data environments.
Apache Flink is an open-source stream processing framework providing powerful capabilities for both data streaming and batch processing.
Apache Kafka is an open-source distributed event streaming platform designed for high-performance data pipelines, streaming analytics, and data integration.
Apache Spark is a unified analytics engine designed for large-scale data processing, offering high-level APIs and an optimized engine for general computation graphs.
Apache Superset is a modern, enterprise-ready business intelligence web application used for data exploration and visualization.
Metabase is an open-source business intelligence and embedded analytics tool designed to allow users to query data and create dashboards without requiring SQL knowledge.
Qdrant is a high-performance vector similarity search engine and database written in Rust, designed for AI applications requiring semantic search and large-scale vector management.
PostgreSQL is an advanced object-relational database management system that supports an extended subset of the SQL standard.
ClickHouse is an open-source, column-oriented database management system designed for real-time analytical data reporting.
DuckDB is a high-performance, in-process analytical SQL database management system designed for speed, portability, and ease of use.