Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Udemy

PySpark Data Engineering: The Complete Real-World Workflow

via Udemy

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Learn PySpark through a real project workflow: interactive dev, modular ETL, Dev/QA, Airflow, Git & production handoff.

What you'll learn:
  • Understand Spark fundamentals and develop PySpark code interactively using Jupyter and PySpark Shell.
  • Turn interactive PySpark code into modular, reusable ETL applications using DataFrames and Spark SQL.
  • Run the same PySpark application across Dev and QA using environment-specific configurations and parameters.
  • Schedule and orchestrate PySpark pipelines using Airflow and manage code through real-world Git workflows.
  • Follow the complete project delivery workflow from development and testing to Jira-based production handoff.

Want to learn PySpark for Data Engineering and understand how it is actually used in a real project?

This course goes beyond isolated PySpark syntax and transformations to show you the complete workflow around PySpark code in a Data Engineering project.

You'll see how PySpark code is developed interactively, turned into modular ETL applications, run across Dev/QA environments, orchestrated with Airflow, managed through Git, monitored using Spark UI and logs, and finally taken through the Production handoff process.


What You'll Learn

  • Develop PySpark interactively using Jupyter and PySpark Shell

  • Build ETL pipelines using DataFrames and Spark SQL

  • Turn interactive code into modular, reusable PySpark applications

  • Structure applications using scripts, configs, environment files, and reusable modules

  • Run applications using spark-submit

  • Run the same application across Dev and QA environments

  • Schedule and orchestrate pipelines using cron and Airflow

  • Manage code using Git branching and merging workflows

  • Monitor and troubleshoot applications using Spark UI and logs

  • Follow the Production handoff and deployment process


Hands-On Project Workflow

You'll work with Spark/PySpark, Airflow, Docker, HDFS, Jupyter, and Git while building end-to-end Data Engineering projects.

The course follows the journey:

Interactive Development → ETL → Modular Application → Dev/QA → Orchestration → Git → Monitoring → Production Handoff


Who Is This For?

This course is for Data Engineers, developers, and ETL professionals who want practical PySpark experience and want to understand how PySpark code fits into the broader Data Engineering project workflow.


By the End

You won't just know how to write PySpark code. You'll understand how that code is developed, structured, executed, orchestrated, managed across environments, monitored, and taken through the Production handoff in a Data Engineering project.

Taught by

Chandra Venkat

Reviews

4.5 rating at Udemy based on 99 ratings

Start your review of PySpark Data Engineering: The Complete Real-World Workflow

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.