Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

Training Vision Language Models from Scratch Using Text-Only LLMs

Neural Breakdown with AVB via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Learn to build Vision Language Models from scratch by transforming text-only Large Language Models into multimodal systems capable of processing both text and images. Explore the Query Former (Q-Former) architecture from the BLIP-2 paper through visual explanations and hands-on coding implementation. Master Vision Transformers and their integration with language models, understand cross-attention mechanisms in transformer architectures, and implement Q-Former models using BERT as a foundation. Follow a comprehensive step-by-step coding guide that covers ViT implementation, Q-Former development, and LORA fine-tuning techniques for language models, providing you with practical skills to train your own Vision Language Models with complete code examples and thorough explanations of multimodal AI architecture.

Syllabus

- Intro
- Vision Transformers
- Coding ViT
- Q-Former models
- Coding Q-Former from a BERT
- Cross Attention in Transformers
- Coding Q-Formers
- LORA finetune Language Model
- Summary

Taught by

Neural Breakdown with AVB

Reviews

5.0 rating, based on 1 Class Central review

Start your review of Training Vision Language Models from Scratch Using Text-Only LLMs

  • Profile image for Farhan Nabi
    Farhan Nabi
    Outstanding
    This course provides an exceptional, hands-on breakdown of how to build and train vision language models completely from scratch using text-only LLMs. The instructor (AVB) does a phenomenal job of explaining complex multimodal architectures, bridging the gap between computer vision and natural language processing in a way that is highly digestible.
    The step-by-step implementation details and practical insights into aligning image embeddings with text spaces are incredibly valuable for anyone looking to understand the underlying mechanics of modern VLMs. Highly recommended for machine learning practitioners and researchers alike!

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.