Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

IBM

Foundations of AI Native Data Engineering

IBM via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
This course introduces the foundational shift from traditional data engineering to AI-native data engineering. Learners reframe data platforms and pipelines as intelligent, production-grade assets that power analytics, machine learning, generative AI systems, semantic search, RAG workflows, and AI-powered assistants. This course explains how AI workloads consume data differently, introduces embeddings and vector retrieval as core primitives, demystifies LLMs as technical systems, and helps learners connect familiar data engineering skills to modern AI data lifecycles and AI-native architectures. The updated Course 1 design also emphasizes that production AI-native systems require ownership boundaries and collaboration across data engineering, ML/AI engineering, application engineering, platform engineering, security/governance, and product teams. By the end of the course, learners understand how modern AI systems are built end to end, how data engineering enables them, and how to redesign legacy pipelines to support semantic search, RAG, and AI-powered workflows.

Syllabus

  • Welcome to the Course
    • Get oriented to Foundations of AI Native Data Engineering and learn how the course is structured, what to expect, and how to prepare for success. You’ll review the learning journey, assessments, final project arc, prerequisites, and responsible use of AI assistants before starting the technical modules.
  • Module 1: Introduction to AI-Native Data Engineering
    • Learn how AI workloads are changing data engineering and why traditional BI-focused pipelines often need redesign to support semantic search, RAG, LLM applications, embeddings, and agentic consumers. You’ll compare ETL and AI-native pipelines, identify architecture and governance gaps, and begin a course project by defining an AI use case, mapping requirements to platform components, and clarifying ownership boundaries.
  • Module 2: AI Fundamentals and Stack Components for Data Engineers
    • This module introduces core AI concepts from a data engineering perspective, helping you distinguish AI, machine learning, deep learning, generative AI, and LLM-based systems. You will also learn how key AI stack components such as data preparation, embeddings, vector databases, retrieval systems, RAG, and serving interfaces work together in real applications.
  • Module 3: The AI Data Lifecycle and the Role of AI Data Engineers
    • Learn how the AI data lifecycle extends data engineering beyond traditional analytics into feature engineering, embeddings, retrieval, inference, evaluation, monitoring, and governance. You’ll examine how AI data engineers design AI-ready pipelines, identify drift and leakage risks, define monitoring signals, and coordinate ownership across teams to support reliable, explainable, and governed AI systems.
  • Module 4: LLMs for Data Engineers: Prompting, Hosting, Fine-Tuning, and Scale
    • Learn how large language models shape data engineering decisions around prompting, context management, hosting, cost, latency, scaling, fine-tuning, and retrieval-augmented generation. You’ll focus on practical architecture, governance, and production tradeoffs rather than model research, and create project artifacts to support LLM integration decisions.
  • Module 5: RAG Pipelines: Design, Implementation, Failure Diagnosis, and Fine-Tuning Tradeoffs
    • Learn how retrieval-augmented generation (RAG) pipelines work and how data engineers design, support, and troubleshoot them in AI-native systems. You will examine the full RAG workflow—from source preparation and chunking to retrieval, prompt assembly, evaluation, governance, and failure diagnosis—and compare when RAG, fine-tuning, or both are the right choice.
  • Module 6: Build a Semantic Search Engine with AI-Native Data Architecture
    • Apply the concepts from Course 1 to design or build a semantic search workflow using AI-native data architecture patterns. You will work with corpus preparation, embeddings, vector retrieval, metadata, governance, and optional RAG-style context assembly, then package your work into a final project-ready semantic search or RAG design.
  • Course Summary
    • Wrap up Course 1 by connecting what you built to portfolio-ready evidence, professional practice, and the next step in the certificate. In this short bookend module, you’ll review your final project artifacts, reinforce the AI-native data engineering mindset, and create a practical plan for refining your work and preparing for Course 2.
  • Final Exam

Taught by

Antonio Cangiano and Ruslan Podgaets

Reviews

Start your review of Foundations of AI Native Data Engineering

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.