Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

IBM

Lakehouse Architecture for AI-Native Data Platforms

IBM via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
This course focuses on designing governed, scalable lakehouse architectures that support AI-native data platforms. Learners translate AI workload requirements into data product SLOs, compare open table formats, design ingestion and replay strategies, manage schema evolution, support reproducibility, and define observability signals for freshness, latency, throughput, and cost. The course emphasizes architecture and operational patterns rather than vendor-specific platform administration. By the end of the course, learners can explain when a lakehouse is preferable to a warehouse or data lake for AI workloads, compare Delta Lake, Iceberg, and Hudi at a practical level, design replay-safe ingestion patterns, and apply governance controls such as data contracts, lineage, access control, audit logging, and CI gates. The result is a production-oriented platform design for governed AI data systems.

Syllabus

  • Welcome to the Course
    • This welcome module introduces Course 4, Lakehouse Architecture for AI-Native Data Platforms, and orients learners to the course purpose, outcomes, and expectations. Learners will see why governed Lakehouse architecture matters for AI-native data engineering, review the major goals of the course, and confirm the prerequisite knowledge needed to succeed.
  • Module 1: Requirements, Data Product SLOs, and Architecture Tradeoffs
    • This module teaches learners to turn AI workload needs into measurable data product SLOs and use those SLOs to evaluate lakehouse, warehouse, and data lake architecture choices. Learners build project-ready requirements, tradeoff notes, and an initial architecture recommendation for AI-native and multimodal data products.
  • Module 2: Table Formats, Time Travel, and Platform Tradeoffs
    • This module teaches learners how to compare Delta Lake, Apache Iceberg, and Apache Hudi for AI native lakehouse workloads, with emphasis on reliability, reproducibility, and governance. Learners design decision artifacts for table format selection, time travel and snapshot strategy, rollback readiness, and catalog integration to support final project architecture choices.
  • Module 3: Schema Evolution, Retention, and Deletion Workflows
    • This module teaches learners how to manage schema evolution, retention, lifecycle states, and deletion workflows in governed Lakehouse environments that support AI workloads. Learners examine compatibility risks across structured, nested, derived, and vectorized data, then design project-ready policies and workflows that preserve reliability, auditability, reproducibility, and compliance readiness.
  • Module 4: Ingestion Architecture: Batch, CDC, Streaming, and Replay
    • Learn how to design reliable Lakehouse ingestion architectures by choosing among batch, CDC, streaming-style, and event-driven patterns based on freshness, latency, cost, and AI workload needs. You will also plan replay, backfill, idempotency, late-data handling, and validation gates so data can be safely promoted to trusted AI-ready states.
  • Module 5: Provisioning, Orchestration, Observability, and Cost SLOs
    • This module teaches learners how to operate an AI-native lakehouse reliably by making concrete decisions about compute provisioning, orchestration, observability, alerting, cost SLOs, and incident response. By the end, learners produce an operations package including a dashboard specification, alert matrix, cost SLOs, and runbook draft for common lakehouse incidents.
  • Module 6: Governance, Security, CI Gates, and Course Project
    • This module brings together governance, security, auditability, and CI controls to finalize a governed Lakehouse architecture for AI workloads. Learners define data contracts and evidence, map lineage and access controls, design CI gates, and complete a portfolio-ready final project package with clear tradeoffs and operational readiness.
  • Course Summary
    • This closing module wraps up Course 4 by celebrating your progress, reflecting on the course’s major architectural themes, and previewing the transition to Course 5. It reinforces how Lakehouse architecture foundations prepare you for unstructured data engineering in AI-native platforms.
  • Final Exam
    • This Final Exam assesses your ability to apply lakehouse architecture concepts across the full course in realistic decision making contexts. You will evaluate tradeoffs, connect requirements to design and operating choices, and demonstrate defensible reasoning for both platform architecture and governance.

Taught by

Antonio Cangiano and Ruslan Podgaets

Reviews

Start your review of Lakehouse Architecture for AI-Native Data Platforms

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.