Training Vision Language Models from Scratch Using Text-Only LLMs
Neural Breakdown with AVB via YouTube
-
12
-
- Write review
AI, Data Science & Cloud Certificates from Google, IBM & Meta
Learn Excel and Financial Modeling the Way Finance Teams Actually Use Them
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Learn to build Vision Language Models from scratch by transforming text-only Large Language Models into multimodal systems capable of processing both text and images. Explore the Query Former (Q-Former) architecture from the BLIP-2 paper through visual explanations and hands-on coding implementation. Master Vision Transformers and their integration with language models, understand cross-attention mechanisms in transformer architectures, and implement Q-Former models using BERT as a foundation. Follow a comprehensive step-by-step coding guide that covers ViT implementation, Q-Former development, and LORA fine-tuning techniques for language models, providing you with practical skills to train your own Vision Language Models with complete code examples and thorough explanations of multimodal AI architecture.
Syllabus
- Intro
- Vision Transformers
- Coding ViT
- Q-Former models
- Coding Q-Former from a BERT
- Cross Attention in Transformers
- Coding Q-Formers
- LORA finetune Language Model
- Summary
Taught by
Neural Breakdown with AVB
Reviews
5.0 rating, based on 1 Class Central review
Showing Class Central Sort
-
Outstanding
This course provides an exceptional, hands-on breakdown of how to build and train vision language models completely from scratch using text-only LLMs. The instructor (AVB) does a phenomenal job of explaining complex multimodal architectures, bridging the gap between computer vision and natural language processing in a way that is highly digestible.
The step-by-step implementation details and practical insights into aligning image embeddings with text spaces are incredibly valuable for anyone looking to understand the underlying mechanics of modern VLMs. Highly recommended for machine learning practitioners and researchers alike!