About this project
This repository is a comprehensive curated list of data engineering tools and resources, organized into clear sections. It covers a wide range of topics including relational, key-value, columnar, document, graph, distributed, and time-series databases; data comparison tools; data ingestion frameworks such as Kafka, Airbyte, Meltano, and dlt; file systems including HDFS, S3, Alluxio, and JuiceFS; serialization formats like Avro, Parquet, ORC, and Protobuf; stream and batch processing tools; charts and dashboards; workflow orchestration; data lake management; the ELK stack; Docker; datasets; monitoring with Prometheus; data profiling; schema management; testing; and community resources including forums, conferences, podcasts, and books. Each entry includes a brief description and a link to the project, making it a practical reference for developers building data pipelines and managing data infrastructure.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.