Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Udemy

Databricks Data Engineering with AWS

via Udemy

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Build a production lakehouse & deploy a real capstone project with Unity Catalog, Delta Lake, Lakeflow, DABs & CI/CD

What you'll learn:
  • Set up and govern a production Databricks workspace on AWS using Unity Catalog
  • Master Delta Lake internals — ACID transactions, time travel, constraints, and performance tuning
  • Design Medallion Architecture pipelines, first manually, then declaratively with Lakeflow Declarative Pipelines
  • Ingest data at scale with Lakeflow Connect — SaaS connectors, database CDC, and Auto Loader
  • Orchestrate production pipelines with Lakeflow Jobs — DAGs, retries, control flow, REST API and CLI
  • Build a complete production lakehouse for a real e-commerce business, from ingestion through five gold-layer outputs
  • Write unit and integration tests for Databricks pipeline code with pytest
  • Package and deploy pipelines using Databricks Asset Bundles (DABs)
  • Build a CI/CD pipeline with GitHub Actions that tests, validates, and deploys to a UAT environment

Databricks has become the default lakehouse platform for data engineering on AWS — over 60% of the Fortune 500 run on it. But knowing individual features isn't the same as being able to design, build, test, and deploy a real production pipeline. This course takes you through both: you'll master the core Databricks and AWS skills chapter by chapter, then apply every one of them to a single, realistic capstone project — an end-to-end lakehouse built for a real business, deployed the way production teams actually deploy.

Learn Databricks on AWS and Build a Real Production Lakehouse from the Ground Up

  • Set up and govern a Databricks workspace on AWS with Unity Catalog from day one

  • Master Delta Lake — ACID transactions, time travel, constraints, performance

  • Design Medallion Architecture pipelines with Lakeflow Connect and Lakeflow Declarative Pipelines

  • Orchestrate production pipelines with Lakeflow Jobs — multi-task DAGs, retries, parameterization

  • Build a complete lakehouse for a real e-commerce business — five source systems, five business-critical gold outputs

  • Test, package, and deploy your pipelines with pytest, Databricks Asset Bundles, and GitHub Actions CI/CD

A complete path from Databricks fundamentals to a deployed, production-grade lakehouse — built one real skill at a time.

Phase 1 — Foundations. You'll start with the core skills every Databricks data engineer needs on AWS:

  • Workspace setup and Unity Catalog governance

  • Delta Lake internals — ACID transactions, time travel, constraints

  • Medallion Architecture, built by hand first, then declaratively with Lakeflow Declarative Pipelines

  • Ingestion with Lakeflow Connect — SaaS, database CDC, and Auto Loader

  • Orchestration with Lakeflow Jobs — DAGs, retries, control flow, REST API and CLI

Phase 2 — The Capstone. Every skill above gets applied to one continuous project: StepRight, a mid-size online footwear retailer with five source systems feeding five gold-layer outputs — daily revenue, customer 360, product performance, funnel analysis, and fulfillment health.

You'll ingest CDC and file-based data at production scale, then go further than most courses do:

  • Write unit and integration tests for your transformation logic

  • Package the project as a Databricks Asset Bundle

  • Wire up GitHub Actions CI/CD — test, validate, deploy to UAT

This is the same workflow real data platform teams run — not a toy example.

By the end of this course, you'll have built and deployed a governed, tested, production-structured lakehouse — end to end, on your own.

You'll walk away with:

  • A complete, working lakehouse project you built and can show, not just watched

  • Hands-on notebooks for every chapter, ready to import into your own Databricks workspace

  • A full GitHub repo structure from the capstone, showing exactly how a production project is organized

This isn't a features tour. It's the architecture, tooling, and deployment discipline real data platform teams run.

Disclaimer: This course was developed with the assistance of AI tools for content research, editing, and slide production. All technical content has been reviewed, tested and validated by the instructor.

Taught by

Prashant Kumar Pandey

Reviews

4.8 rating at Udemy based on 50 ratings

Start your review of Databricks Data Engineering with AWS

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.