Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Video RAG Foundations

Edureka via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course introduces retrieval augmented generation and extends it from text documents to video. You build a working document-based RAG chatbot first, then learn to break video into the signals a retrieval system can search. You start with RAG architecture, embeddings, and vector stores, building a chatbot that answers questions from your own documents. You then process raw video, extracting frames, audio, transcripts, captions, and on-screen text, and combine those signals into structured, timestamped records. The course closes with multimodal embeddings, vector storage in ChromaDB, and your first natural-language video search application. By the end of this course, you will be able to: 1. Explain RAG architecture and the role of embeddings, retrievers, and vector stores. 2. Build a document-based RAG chatbot using LangChain and a vector database. 3. Extract frames, audio, transcripts, captions, and OCR text from raw video. 4. Combine multimodal signals into structured, timestamped video records. 5. Generate multimodal embeddings and store them for semantic retrieval. 6. Build a search application that returns video segments from a plain-language query. This course is intended for Python developers, data engineers, and AI practitioners. You should be comfortable writing Python and running notebooks. Turn raw video into searchable records you can query in plain language.

Syllabus

  • Introduction to VideoRAG
    • This module introduces the foundations of Retrieval-Augmented Generation and explores how VideoRAG extends traditional RAG workflows to process and retrieve information from video content. Learners explore VideoRAG architecture, key components, and supporting frameworks used to connect visual, audio, and textual information for more contextual and grounded AI responses.
  • Preparing Videos for Retrieval
    • This module focuses on preparing raw video content for retrieval by transforming unstructured video data into structured and searchable information. Learners explore video processing workflows, including frame extraction, transcription, scene segmentation, chunking strategies, and metadata generation to create effective video knowledge sources for VideoRAG systems.
  • Embeddings and Semantic Search
    • This module introduces embeddings, vector databases, and semantic search techniques that enable efficient retrieval in VideoRAG systems. Learners explore how video content is converted into meaningful representations, stored, and searched using similarity-based retrieval approaches to identify relevant information and support accurate AI-generated responses.

Taught by

Edureka

Reviews

Start your review of Video RAG Foundations

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.