Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Most AI systems still expect clean text. The work people need done arrives as screenshots, scanned invoices, product photos, voice notes, and charts buried in PDFs. This Specialization covers multimodal AI agents: systems that interpret all of it, then act on what they find.
You start with what separates an agent from an assistant, then add prompt design, memory, and tool calling until it does something useful. You bring in vision-language models for multimodal search, coordinate agents on one task, then test, guard, and ship what you built.
By the end of this Specialization, you will be able to:
1. Describe the agent lifecycle from perception through reasoning to action.
2. Process text, image, audio, video, and document input.
3. Implement tool calling with reliable schemas and error handling.
4. Build multimodal retrieval over documents containing charts and tables.
5. Coordinate multiple agents through defined roles and shared context.
6. Test, guard, and deploy agents with REST APIs and containers.
This Specialization suits AI engineers, machine learning engineers, applied AI developers, automation engineers, and solution architects moving past prompt-based chatbots into systems that act. It assumes basic Python and comfort calling APIs, and no background in agents, computer vision, or speech processing.
Enroll now to build multimodal agents that see, listen, reason, and act.
Syllabus
- Course 1: Multimodal AI and Agent Fundamentals
- Course 2: Multimodal Agents with Vision Language Models
- Course 3: Multimodal Agent Testing and Deployment
Courses
-
This course covers testing, guardrails, and deployment for multimodal agents. Evaluating open-ended behavior and operating agents safely are the two areas where applied agent projects most often fail. You start by writing tests for agent responses and building evaluation criteria for behavior that has no single correct answer. You then implement input and output guardrails, log decision traces so agent choices can be audited, and apply responsible AI practices covering bias, transparency, and disclosure. The course closes with deployment: Streamlit interfaces, REST APIs, and containers, alongside the operating concerns of latency, token cost, and rate limiting, applied through builds including support ticket triage, visual search, and a voice-driven assistant. By the end of this course, you will be able to: - Write tests and evaluation criteria for open-ended agent responses. - Implement input and output guardrails that block unsafe or invalid behavior. - Log and audit agent decision traces. - Apply responsible AI practices for bias, transparency, and disclosure. - Deploy agents using Streamlit, REST APIs, and containers. - Measure and manage latency, token cost, and rate limits. Intended for learners who have completed the earlier courses in the Specialization. Enroll now to make your agents testable, safe, and ready for other people to use.
-
This course covers the fundamentals of multimodal AI agents: what separates an agent from an assistant, and how agents handle text, image, audio, video, and document input. These foundations carry through the Specialization. You start with how large language models work, covering tokens, context windows, and their limitations, then examine the agent lifecycle from perception through reasoning to action. You then process each modality in turn: generating and structuring text, captioning and reasoning over images, transcribing audio, summarizing video, and extracting tables from PDF documents. The course closes by combining prompt design, context handling, memory, and tool calling into your first working multimodal agent. By the end of this course, you will be able to: - Distinguish AI assistants from AI agents and describe the agent lifecycle. - Explain how LLMs process tokens and context, and where they fail. - Process text, image, audio, video, and document inputs with multimodal models. - Design prompts and context handling that keep an agent coherent. - Implement short-term and long-term memory in an agent. - Build an agent that combines multimodal input with a tool call. Intended for Python developers, AI engineers, and data scientists. You should be able to write basic Python and work with APIs. Enroll now to build your first agent that handles more than plain text.
-
This course covers vision-language models and multi-agent coordination: how agents interpret images alongside text and divide work between them. Together they take an agent beyond single inputs. You explore how vision-language models such as CLIP, BLIP, and LLaVA align images with text, and use those embeddings to build a multimodal search system. You then construct a multimodal RAG pipeline that answers questions about charts and tables inside PDF documents, where text-only retrieval fails. The course closes with orchestration: planner, executor, and critic patterns that break complex tasks into steps, reliable tool schemas with error handling, and collaborative multi-agent systems where specialized agents pass context between each other. By the end of this course, you will be able to: - Explain how vision-language models align image and text representations. - Build a multimodal search system using shared embedding spaces. - Construct a multimodal RAG pipeline over documents containing charts. - Apply planner, executor, and critic patterns to decompose complex tasks. - Implement tool calling with reliable schemas and error handling. - Coordinate multiple agents through defined roles, handoffs, and shared context. Intended for learners who have completed Multimodal AI and Agent Fundamentals. Enroll now to give your agents sight, retrieval, and the ability to work as a team.
Taught by
Edureka