OPEN SOURCE, OPEN TO EVERYONE

Open source. Open possibilities.

Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.

Human-curated · Discover open source4101discovered

A little curiosity. A world of open source.

THE FIRST COLLECTION
Topic: spark清除
deltadelta-io
ADDED

Delta Lake is an open-source storage framework enabling Lakehouse architecture. It integrates with compute engines like Spark, PrestoDB, Flink, Trino, and Hive, offering ACID transactions, time travel, and schema evolution for reliable data management.

Data & databasesStreaming & warehousing
redashgetredash
ADDED

Redash is a browser-based data exploration and visualization tool that allows users to query, visualize, and share data from various SQL and NoSQL sources.

Data & databasesProductivityAnalytics & BI
sqlglottobymao
ADDED

SQLGlot is a pure‑Python SQL parser, transpiler, optimizer, and engine that supports over 30 dialects, enabling formatting, cross‑dialect translation, AST manipulation, and optional C‑extension acceleration for fast SQL processing.

Developer toolsData & databasesComponent libraries
deeplearning4jdeeplearning4j
ADDED

Eclipse Deeplearning4J (DL4J) is a JVM-based deep learning ecosystem supporting Java, Scala, Kotlin, and more. It includes model import from Keras, TensorFlow, ONNX; ND4J for linear algebra; SameDiff for automatic differentiation; DataVec for ETL; and runs on CPU/GPU across Windows, Linux, macOS.

AI & MLDeveloper toolsTraining & evaluation

A runnable local real-time analytics pipeline using Kafka-compatible broker (Redpanda) and PySpark Structured Streaming. Simulates clickstream events, performs windowed aggregation with watermarking, all via docker-compose.

Developer toolsAutomationBackend & APIs
sparkapache
ADDED

Apache Spark is a unified analytics engine designed for large-scale data processing, offering high-level APIs and an optimized engine for general computation graphs.

Data & databasesData integration