About this project
Data Engineering Zoomcamp is a free, open-source 9-week course from DataTalks.Club that teaches data engineering fundamentals by having learners build an end-to-end data pipeline from scratch. Materials, video lectures and homework are freely available in the repository.
The course targets developers, analysts and data scientists who want to learn how to build data pipelines and work with a modern data engineering stack. Prerequisites are basic coding experience and familiarity with SQL; Python is helpful but not required, and no prior data engineering experience is needed.
It can be followed in two ways. A live cohort runs with deadlines, graded homework, a leaderboard, peer review and certificate eligibility; lectures are pre-recorded, so "live" refers to the cohort structure rather than live classes. Self-paced learners can start anytime and use the same materials, with homework available but not scored and no certificate.
The syllabus is organized into modules and workshops:
- Module 1: Containerization and Infrastructure as Code — introduction to GCP, Docker and Docker Compose, running PostgreSQL with Docker, and infrastructure setup with Terraform.
- Module 2: Workflow Orchestration — data lakes and workflow orchestration, with Kestra.
- Workshop 1: Data Ingestion — API reading, pipeline scalability, data normalization and incremental loading, using dlt.
- Module 3: Data Warehousing — introduction to BigQuery, partitioning, clustering and best practices, plus machine learning in BigQuery.
- Module 4: Analytics Engineering — analytics engineering and data modeling, dbt with DuckDB and BigQuery, testing, documentation and deployment.
- Module 5: Data Platforms — building end-to-end pipelines with Bruin, covering ingestion, transformation and quality, and deployment to cloud (BigQuery).
- Module 6: Batch Processing — Apache Spark, DataFrames and SQL, and internals of GroupBy and joins.
- Module 7: Streaming — Kafka, Kafka Streams and KSQL, and schema management with Avro.
A final project applies the concepts in a real-world scenario and includes a peer review and feedback process. Certificates are awarded to learners who complete the final project during a live cohort.
The repository also links to a YouTube lecture playlist, course documentation, a course platform for deadlines and homework, a Slack channel for discussion and troubleshooting, Telegram announcements, and an FAQ. Instructors listed include Alexey Grigorev, Michael Shoemaker, Will Russell, Anna Geller, Juan Manuel Perafan and Arsalan Noorafkan, with several past instructors also credited. Course sponsors include Kestra, Bruin and dltHub. The project is part of DataTalks.Club, a global online data community that organizes events and free courses.
Comments
0 Rating appears after 10 ratings
Sign in to join the discussion.