Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

IBM

Vector Databases and Retrieval Data Engineering

IBM via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
This course teaches learners to design, operate, secure, and evaluate vector-based retrieval systems used in semantic search and RAG applications. Learners work with embeddings, vector schemas, index design, refresh strategies, consistency checks, evaluation signals, retrieval observability, and permissions-aware access controls. The course focuses on practical system design and tradeoffs rather than low-level algorithm implementation. By the end of the course, learners can explain how embeddings enable semantic retrieval, compare dense, sparse, and hybrid retrieval patterns, design vector database schemas and refresh workflows, evaluate retrieval quality, and apply governance and security controls to retrieval systems. Topics include vector stores, metadata filtering, HNSW and IVF concepts, recall/latency tradeoffs, retrieval drift, and audit logging.

Syllabus

  • Module Purpose
    • This welcome module introduces Course 3, explains its role in the AI-Native Data Engineering Professional Certificate, and orients learners to the course journey ahead. Learners review the course purpose, major learning goals, recommended prerequisites, and where to find the full syllabus.
  • Module 1: Embeddings and Vector Retrieval Foundations
    • This module introduces how embeddings power semantic retrieval and how vector retrieval differs from keyword search in practical data engineering workflows. Learners compare similarity metrics and retrieval patterns, then design an initial metadata-rich vector schema and retrieval use case that will serve as the foundation for later vector database design work.
  • Module 2: Vector Database Design and Storage Patterns
    • This module teaches learners how to design vector database storage patterns that connect embeddings, metadata, filters, source links, and source-of-truth systems. Learners build practical, project-ready artifacts for schema design, filtering strategy, integration boundaries, and storage tradeoff documentation in retrieval and RAG architectures.
  • Module 3: Index Design and Scaling
    • This module teaches you how to choose and justify vector index strategies by comparing exact and approximate retrieval, practical ANN structures, compression options, and scaling patterns. You will benchmark tradeoffs across recall, latency, memory, build time, and cost, then turn that evidence into a project ready index strategy and scaling memo.
  • Module 4: Refresh, Updates, Deletes, and Consistency
    • Learn how to keep vector indexes accurate and safe as source data changes by designing refresh, update, delete, rollback, and validation workflows. The module emphasizes operational reliability, governance, and auditability so retrieval systems stay consistent with source-of-truth data over time.
  • Module 5: Retrieval Evaluation and Observability
    • Learn how to evaluate retrieval systems with golden query sets, relevance labels, and core retrieval metrics, then extend that evidence into production observability with latency, throughput, drift, dashboards, and alerts. The module emphasizes turning metric patterns into actionable engineering decisions, documented limitations, and improvement plans.
  • Module 6: Security, Governance, and Course Project
    • This module focuses on securing and governing vector retrieval systems while preparing the final course project package. Learners design access controls, tenant isolation, permissions-aware retrieval, auditability, and governance evidence, then assemble and present a complete engineering handoff for a secure vector retrieval subsystem.
  • Course Summary
    • This closing module helps learners reflect on the broad capabilities they developed in Course 3 and recognize their growth in thinking about vector retrieval as an engineered, governable, and operable subsystem. It also provides a high-level transition to Course 4 by previewing the shift from retrieval systems to Lakehouse architecture for AI-native data platforms.
  • Final Exam
    • The Final Exam assesses your ability to apply vector database and retrieval engineering concepts across realistic production scenarios. It combines a quiz on architecture, operations, evaluation, and security tradeoffs with a case study focused on end to end retrieval system design decisions.

Taught by

Antonio Cangiano and Ruslan Podgaets

Reviews

Start your review of Vector Databases and Retrieval Data Engineering

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.