Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This Specialization covers practical data engineering using widely adopted open-source technologies.You will begin with relational data modeling, SQL, file processing, and data quality using Python and PostgreSQL. You will then process data at scale with PySpark and Spark SQL, build layered and tested transformation models with dbt Core, orchestrate batch workflows with Apache Airflow, and manage open lakehouse data using Apache Iceberg and MinIO.
You will also build streaming pipelines with Apache Kafka and Spark Structured Streaming, applying event-time processing, windowing, checkpoints, monitoring, and recovery.By the end of this Specialization, you will be able to:
Build and validate relational data workflows. Process large datasets with PySpark and Spark SQL. Create layered, tested dbt models. Orchestrate batch pipelines with Apache Airflow. Manage open lakehouse data with Apache Iceberg and MinIO. Build and monitor streaming pipelines with Kafka and Spark Structured Streaming.
It is Designed for aspiring data engineers, analytics engineers, software developers, analysts, and database professionals. Basic Python and SQL knowledge is recommended; prior Spark or Kafka experience is not required.
Enroll now to master open-source data engineering across relational workflows, scalable lakehouse pipelines, and real-time streaming with SQL, Python, Spark, dbt, Airflow, Iceberg, and Kafka.
Syllabus
- Course 1: Data Engineering Foundations with SQL and Python
- Course 2: Lakehouse Data Pipelines with Spark and dbt
- Course 3: Streaming Data Pipelines with Kafka and Spark
Courses
-
Build a strong practical foundation in modern data engineering using Python, PostgreSQL, SQL, Git, GitHub, and common engineering data formats. You will learn how data engineers organize development environments, work with structured data, design relational models, build reliable file-to-database workflows, and prepare trusted datasets for analysis. You will begin by exploring the role of data engineering in modern data systems, the data engineering lifecycle, and the differences between batch and real-time processing. You will also examine important design considerations such as latency, throughput, scalability, and system boundaries. Through guided demonstrations, you will set up a data engineering workspace, work with the command line, organize project files, and manage code using Git and GitHub. You will then work with CSV, JSON, and Parquet datasets and explore concepts such as schemas, data contracts, schema drift, storage choices, relational databases, and data modeling. Using PostgreSQL and SQL, you will create tables, explore data, perform joins, calculate metrics, and use Common Table Expressions and window functions for more advanced analysis. You will also apply defensive SQL practices for handling NULL values, data types, and inconsistent records. Finally, you will build reliable data workflows by loading file-based data into PostgreSQL, validating data quality, cleaning and preparing datasets, and applying principles such as idempotency, incremental loading, safe reprocessing, and troubleshooting. The course concludes with a structured data processing project that brings together ingestion, storage, transformation, validation, and workflow reliability. By the end of this course, you will be able to: - Explain the role of data engineering in modern data systems. - Describe the stages of the data engineering lifecycle. - Differentiate between batch and real-time processing. - Set up and manage a reproducible data engineering workspace. - Work with CSV, JSON, and Parquet data formats. - Understand schemas, data contracts, schema drift, and storage choices. - Design and query relational data using PostgreSQL and SQL. - Use joins, aggregations, CTEs, and window functions for data analysis. - Apply defensive SQL techniques for NULL values and data types. - Build reliable file-to-database workflows. - Validate, clean, and prepare data for downstream analysis. - Apply idempotency, incremental loading, safe reprocessing, and troubleshooting practices. Designed for aspiring data engineers, data analysts, software developers, database professionals, and students entering the data field, this course prepares you to build structured, reliable, and maintainable data workflows using widely used open-source technologies.
-
Build practical expertise in modern batch data engineering using Apache Spark, PySpark, dbt Core, Apache Airflow, Apache Iceberg, MinIO, PostgreSQL, and SQL. You will learn how data engineers process large datasets, build reusable transformation models, orchestrate dependent workflows, and manage analytical data using open lakehouse technologies. You will begin by exploring distributed data processing fundamentals and the architecture of Apache Spark. You will examine drivers, executors, jobs, stages, tasks, lazy evaluation, partitions, shuffles, and execution plans. Through guided demonstrations, you will set up PySpark locally, read CSV, JSON, and Parquet datasets, define schemas, transform DataFrames, and analyse data using Spark SQL. You will then move into transformation engineering with dbt Core, where you will explore ETL and ELT approaches, layered modelling, model grain, and reusable SQL transformations. Using dbt Core with PostgreSQL, you will build staging and intermediate models, create fact and dimension tables, apply tests, generate documentation, inspect lineage, and work with materialization strategies, transformation contracts, and change management. Finally, you will learn how to coordinate batch workflows using Apache Airflow and manage lakehouse data using MinIO and Apache Iceberg. You will work with DAGs, tasks, schedules, dependencies, retries, monitoring, backfills, and catchup. You will also explore object storage, Parquet, Iceberg tables, snapshots, schema evolution, compaction, and small-file management. The course concludes with an orchestrated batch lakehouse project that brings together distributed processing, transformation, orchestration, and open lakehouse storage. By the end of this course, you will be able to: - Explain the fundamentals of distributed data processing and Apache Spark architecture. - Process CSV, JSON, and Parquet datasets using PySpark DataFrames. - Transform, clean, aggregate, and analyse data using PySpark and Spark SQL. - Interpret partitions, shuffles, execution plans, caching, and recomputation. - Explain the role of ETL, ELT, and analytics engineering in modern data workflows. - Build layered transformation models using dbt Core. - Create staging, intermediate, fact, dimension, and mart models. - Apply dbt testing, documentation, lineage, materialization, and transformation contracts. - Orchestrate batch workflows using Apache Airflow. - Manage schedules, dependencies, retries, backfills, catchup, and monitoring. - Create and work with Apache Iceberg tables on MinIO. - Understand snapshots, schema evolution, compaction, and small-file management. - Build an end-to-end orchestrated batch lakehouse pipeline using open-source technologies. Designed for data engineers, aspiring data engineers, analytics engineers, software developers, data analysts, and database professionals, this course prepares you to build scalable, maintainable, and reliable batch data pipelines using modern open-source processing, transformation, orchestration, and lakehouse technologies.
-
Build practical expertise in real-time data engineering using Apache Kafka, Apache Spark, Spark Structured Streaming, PySpark, Python, Docker, and Docker Compose. You will learn how data engineers design, process, monitor, and maintain streaming pipelines that continuously move event data from source systems to reliable downstream outputs. You will begin by exploring event streaming fundamentals and the architecture of Apache Kafka. You will examine brokers, topics, partitions, producers, consumers, offsets, consumer groups, message ordering, and delivery behavior. Through guided demonstrations, you will create Kafka topics, publish events with producers, consume messages, and observe how Kafka distributes and manages streaming data. You will then move into stream processing with Spark Structured Streaming, where you will work with streaming DataFrames, schemas, micro-batch execution, transformations, aggregations, and output sinks. You will also explore event-time processing, late-arriving data, stateful operations, checkpointing, and recovery to understand how Spark maintains progress and processes continuously arriving events reliably. Finally, you will focus on streaming data quality, monitoring, failure handling, and reliable delivery. You will validate streaming records, identify malformed or problematic events, monitor pipeline behavior, troubleshoot processing issues, and apply recovery practices for continuous workloads. The course concludes with a reliable open-source streaming pipeline project that brings together Kafka-based event ingestion, Spark processing, quality validation, monitoring, recovery, and dependable output delivery. By the end of this course, you will be able to: - Explain the fundamentals of event streaming and Apache Kafka architecture. - Work with Kafka brokers, topics, partitions, producers, consumers, offsets, and consumer groups. - Explain how partitioning, ordering, and delivery behavior influence streaming pipelines. - Create and manage Kafka-based producer and consumer workflows. - Process continuously arriving events using Spark Structured Streaming. - Apply schemas and transformations to streaming DataFrames. - Perform streaming aggregations and deliver processed results to output sinks. - Work with event-time processing, late-arriving data, and stateful operations. - Apply checkpointing and recovery techniques to maintain processing continuity. - Validate streaming events and handle malformed or unreliable records. - Monitor Kafka and Spark streaming workflows and identify operational issues. - Troubleshoot common failures across streaming data pipelines. - Apply reliability practices for continuous processing and delivery. - Build an end-to-end streaming data pipeline using Kafka and Spark. Designed for data engineers, aspiring streaming data engineers, software developers, data platform professionals, and technical professionals working with real-time data systems, this course prepares you to build scalable, reliable, and maintainable streaming pipelines using modern open-source technologies.
Taught by
Edureka