Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Multimodal Agents with Vision Language Models

Edureka via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course covers vision-language models and multi-agent coordination: how agents interpret images alongside text and divide work between them. Together they take an agent beyond single inputs. You explore how vision-language models such as CLIP, BLIP, and LLaVA align images with text, and use those embeddings to build a multimodal search system. You then construct a multimodal RAG pipeline that answers questions about charts and tables inside PDF documents, where text-only retrieval fails. The course closes with orchestration: planner, executor, and critic patterns that break complex tasks into steps, reliable tool schemas with error handling, and collaborative multi-agent systems where specialized agents pass context between each other. By the end of this course, you will be able to: - Explain how vision-language models align image and text representations. - Build a multimodal search system using shared embedding spaces. - Construct a multimodal RAG pipeline over documents containing charts. - Apply planner, executor, and critic patterns to decompose complex tasks. - Implement tool calling with reliable schemas and error handling. - Coordinate multiple agents through defined roles, handoffs, and shared context. Intended for learners who have completed Multimodal AI and Agent Fundamentals. Enroll now to give your agents sight, retrieval, and the ability to work as a team.

Syllabus

  • Vision-Language Foundations and Multimodal RAG
    • This module introduces vision-language models, embeddings, and retrieval techniques used to connect visual and textual information. Learners explore multimodal RAG workflows and build systems that retrieve relevant content from images and documents to support more grounded responses.
  • Agent Architectures and Tool Integration
    • This module focuses on how multimodal agents are structured and how they interact with external tools to complete tasks. Learners explore agent design patterns, reliable tool schemas, and tool-calling workflows to build agents for tasks such as web summarisation and image generation.
  • Multi-Agent Collaboration
    • This module focuses on how multiple specialised agents communicate, coordinate, and share responsibilities within a common workflow. Learners explore handoff patterns, context and memory management, and collaborative agent designs to build multimodal multi-agent systems.

Taught by

Edureka

Reviews

Start your review of Multimodal Agents with Vision Language Models

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.