What you'll learn:
- Set up and govern a production Databricks workspace on AWS using Unity Catalog
- Master Delta Lake internals — ACID transactions, time travel, constraints, and performance tuning
- Design Medallion Architecture pipelines, first manually, then declaratively with Lakeflow Declarative Pipelines
- Ingest data at scale with Lakeflow Connect — SaaS connectors, database CDC, and Auto Loader
- Orchestrate production pipelines with Lakeflow Jobs — DAGs, retries, control flow, REST API and CLI
- Build a complete production lakehouse for a real e-commerce business, from ingestion through five gold-layer outputs
- Write unit and integration tests for Databricks pipeline code with pytest
- Package and deploy pipelines using Databricks Asset Bundles (DABs)
- Build a CI/CD pipeline with GitHub Actions that tests, validates, and deploys to a UAT environment
Databricks has become the default lakehouse platform for data engineering on AWS — over 60% of the Fortune 500 run on it. But knowing individual features isn't the same as being able to design, build, test, and deploy a real production pipeline. This course takes you through both: you'll master the core Databricks and AWS skills chapter by chapter, then apply every one of them to a single, realistic capstone project — an end-to-end lakehouse built for a real business, deployed the way production teams actually deploy.
Learn Databricks on AWS and Build a Real Production Lakehouse from the Ground Up
Set up and govern a Databricks workspace on AWS with Unity Catalog from day one
Master Delta Lake — ACID transactions, time travel, constraints, performance
Design Medallion Architecture pipelines with Lakeflow Connect and Lakeflow Declarative Pipelines
Orchestrate production pipelines with Lakeflow Jobs — multi-task DAGs, retries, parameterization
Build a complete lakehouse for a real e-commerce business — five source systems, five business-critical gold outputs
Test, package, and deploy your pipelines with pytest, Databricks Asset Bundles, and GitHub Actions CI/CD
A complete path from Databricks fundamentals to a deployed, production-grade lakehouse — built one real skill at a time.
Phase 1 — Foundations. You'll start with the core skills every Databricks data engineer needs on AWS:
Workspace setup and Unity Catalog governance
Delta Lake internals — ACID transactions, time travel, constraints
Medallion Architecture, built by hand first, then declaratively with Lakeflow Declarative Pipelines
Ingestion with Lakeflow Connect — SaaS, database CDC, and Auto Loader
Orchestration with Lakeflow Jobs — DAGs, retries, control flow, REST API and CLI
Phase 2 — The Capstone. Every skill above gets applied to one continuous project: StepRight, a mid-size online footwear retailer with five source systems feeding five gold-layer outputs — daily revenue, customer 360, product performance, funnel analysis, and fulfillment health.
You'll ingest CDC and file-based data at production scale, then go further than most courses do:
Write unit and integration tests for your transformation logic
Package the project as a Databricks Asset Bundle
Wire up GitHub Actions CI/CD — test, validate, deploy to UAT
This is the same workflow real data platform teams run — not a toy example.
By the end of this course, you'll have built and deployed a governed, tested, production-structured lakehouse — end to end, on your own.
You'll walk away with:
A complete, working lakehouse project you built and can show, not just watched
Hands-on notebooks for every chapter, ready to import into your own Databricks workspace
A full GitHub repo structure from the capstone, showing exactly how a production project is organized
This isn't a features tour. It's the architecture, tooling, and deployment discipline real data platform teams run.
Disclaimer: This course was developed with the assistance of AI tools for content research, editing, and slide production. All technical content has been reviewed, tested and validated by the instructor.