Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Lakehouse Data Pipelines with Spark and dbt

Edureka via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Build practical expertise in modern batch data engineering using Apache Spark, PySpark, dbt Core, Apache Airflow, Apache Iceberg, MinIO, PostgreSQL, and SQL. You will learn how data engineers process large datasets, build reusable transformation models, orchestrate dependent workflows, and manage analytical data using open lakehouse technologies. You will begin by exploring distributed data processing fundamentals and the architecture of Apache Spark. You will examine drivers, executors, jobs, stages, tasks, lazy evaluation, partitions, shuffles, and execution plans. Through guided demonstrations, you will set up PySpark locally, read CSV, JSON, and Parquet datasets, define schemas, transform DataFrames, and analyse data using Spark SQL. You will then move into transformation engineering with dbt Core, where you will explore ETL and ELT approaches, layered modelling, model grain, and reusable SQL transformations. Using dbt Core with PostgreSQL, you will build staging and intermediate models, create fact and dimension tables, apply tests, generate documentation, inspect lineage, and work with materialization strategies, transformation contracts, and change management. Finally, you will learn how to coordinate batch workflows using Apache Airflow and manage lakehouse data using MinIO and Apache Iceberg. You will work with DAGs, tasks, schedules, dependencies, retries, monitoring, backfills, and catchup. You will also explore object storage, Parquet, Iceberg tables, snapshots, schema evolution, compaction, and small-file management. The course concludes with an orchestrated batch lakehouse project that brings together distributed processing, transformation, orchestration, and open lakehouse storage. By the end of this course, you will be able to: - Explain the fundamentals of distributed data processing and Apache Spark architecture. - Process CSV, JSON, and Parquet datasets using PySpark DataFrames. - Transform, clean, aggregate, and analyse data using PySpark and Spark SQL. - Interpret partitions, shuffles, execution plans, caching, and recomputation. - Explain the role of ETL, ELT, and analytics engineering in modern data workflows. - Build layered transformation models using dbt Core. - Create staging, intermediate, fact, dimension, and mart models. - Apply dbt testing, documentation, lineage, materialization, and transformation contracts. - Orchestrate batch workflows using Apache Airflow. - Manage schedules, dependencies, retries, backfills, catchup, and monitoring. - Create and work with Apache Iceberg tables on MinIO. - Understand snapshots, schema evolution, compaction, and small-file management. - Build an end-to-end orchestrated batch lakehouse pipeline using open-source technologies. Designed for data engineers, aspiring data engineers, analytics engineers, software developers, data analysts, and database professionals, this course prepares you to build scalable, maintainable, and reliable batch data pipelines using modern open-source processing, transformation, orchestration, and lakehouse technologies.

Syllabus

  • Distributed Processing with PySpark
    • Explore distributed data processing with Apache Spark and learn to work with PySpark for large-scale data transformation and analysis. Understand Spark’s execution model, partitions, shuffles, and performance optimisation.
  • Transformation Engineering with dbt
    • Learn modern data transformation and modelling using dbt Core. Build reliable data models while applying testing, documentation, and materialisation strategies to improve data quality.
  • Orchestration and Lakehouse Foundations
    • Learn to orchestrate data workflows with Apache Airflow and explore open lakehouse technologies such as Iceberg and MinIO. Build an end-to-end batch pipeline using these open-source tools.

Taught by

Edureka

Reviews

Start your review of Lakehouse Data Pipelines with Spark and dbt

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.