Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Multimodal AI and Agent Fundamentals

Edureka via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course covers the fundamentals of multimodal AI agents: what separates an agent from an assistant, and how agents handle text, image, audio, video, and document input. These foundations carry through the Specialization. You start with how large language models work, covering tokens, context windows, and their limitations, then examine the agent lifecycle from perception through reasoning to action. You then process each modality in turn: generating and structuring text, captioning and reasoning over images, transcribing audio, summarizing video, and extracting tables from PDF documents. The course closes by combining prompt design, context handling, memory, and tool calling into your first working multimodal agent. By the end of this course, you will be able to: - Distinguish AI assistants from AI agents and describe the agent lifecycle. - Explain how LLMs process tokens and context, and where they fail. - Process text, image, audio, video, and document inputs with multimodal models. - Design prompts and context handling that keep an agent coherent. - Implement short-term and long-term memory in an agent. - Build an agent that combines multimodal input with a tool call. Intended for Python developers, AI engineers, and data scientists. You should be able to write basic Python and work with APIs. Enroll now to build your first agent that handles more than plain text.

Syllabus

  • AI, LLM and Agent Foundations
    • This module introduces the foundations of AI, Large Language Models, and AI agents, helping learners understand how these technologies differ and work together. Learners explore prompting, reasoning, tool use, and basic agent workflows to build a strong foundation for multimodal AI.
  • Introduction to Multimodal AI and Modalities
    • This module focuses on how multimodal AI works with text, images, audio, video, and documents within a single system. Learners explore practical workflows for processing different input types and understand how multimodal models combine information across modalities.
  • Building Blocks of Multimodal Agents
    • This module focuses on the core components required to build multimodal agents, including prompts, tools, context, memory, and knowledge retrieval. Learners apply these components to create structured agent workflows and develop a basic multimodal AI agent.

Taught by

Edureka

Reviews

Start your review of Multimodal AI and Agent Fundamentals

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.