Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

Coursera

Transformers for Vision AI, Multimodal & Generative AI

Packt via Coursera

Overview

Google, IBM & Meta Certificates – 40% Off
One plan covers every Professional Certificate on Coursera.
Unlock All Certificates
Discover how transformer architectures are revolutionizing computer vision and multimodal AI, and explore the latest advancements toward general artificial intelligence. Learn to apply transformers to image, video, and cross-modal tasks. This course explores the expanding frontier of transformer models beyond natural language, focusing on their applications in computer vision, multimodal AI, and generative ideation. Learners will investigate vision transformers, text-to-image and text-to-video generation, and the integration of multiple AI models for advanced tasks. The course also addresses risk mitigation in large models and looks ahead to the future of AI with functional AGI and creative generative systems. By the end, you will understand how to harness transformers for cutting-edge vision and multimodal applications. With a focus on emerging trends and practical implementations, this course guides learners through the latest research and real-world use cases in vision and multimodal AI. Concepts are introduced progressively, enabling learners to build expertise in applying transformers across diverse domains. This course is part three of a three-course Specialization designed to build a complete and cohesive understanding of the subject. While it offers valuable skills on its own, you'll gain the most benefit by progressing through all three courses as a structured learning journey. This course is based on Transformers for Natural Language Processing and Computer Vision, by Denis Rothman. Packt is one of the world's most prolific publishers of cutting-edge technical content. For over two decades we've made it our mission to curate and publish the knowledge of only the very best technical experts. We focus on real-world courses that help our customers get the job done, with coverage that extends across a wide range of established and cutting-edge technical topics. If you're an individual or an organisation that embraces learning by doing, Packt is the perfect fit for you.

Syllabus

  • Beyond Text: Vision Transformers in the Dawn of Revolutionary AI
    • This module introduces learners to vision transformers and multimodal AI models, including ViT, CLIP, and DALL-E. You will explore how images are processed as input for transformers, understand the architecture and configuration of feature extractors, and examine real-world applications of these models in creative and mainstream contexts.
  • Transcending the Image-Text Boundary with Stable Diffusion
    • This module introduces the fundamentals of Stable Diffusion for generating images and videos from text prompts. Learners will explore the underlying architecture, practical implementations using Keras and Hugging Face, and the adaptation of OpenAI CLIP for text-to-video synthesis. By the end, participants will gain hands-on experience running diffusion models and understanding their creative potential.
  • Hugging Face AutoTrain: Training Vision Models without Coding
    • This module guides learners through the process of training and deploying vision models using Hugging Face AutoTrain, all without writing code. You will explore data preparation, model selection, and evaluation techniques, while gaining hands-on experience with popular architectures like ViT, BEiT, and ConvNext.
  • On the Road to Functional AGI with HuggingGPT and its Peers
    • This module introduces learners to the concept of Functional Artificial General Intelligence (F-AGI) and demonstrates how advanced AI platforms like HuggingGPT and Google Cloud Vision can be integrated for complex task automation. Learners will explore model chaining, cross-platform AI pipelines, and practical approaches to enhancing computer vision accuracy in challenging scenarios.
  • Beyond Human-Designed Prompts with Generative Ideation
    • This module introduces automated generative ideation systems, demonstrating how AI tools like ChatGPT, Llama 2, Midjourney, and Microsoft Designer can streamline content and image creation without manual prompts. Learners will explore practical workflows, ethical considerations, and integration strategies for efficient, scalable ideation. By the end, you'll be equipped to implement and extend automated pipelines for creative tasks.

Taught by

Packt - Course Instructors

Reviews

Start your review of Transformers for Vision AI, Multimodal & Generative AI

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.