Emerging Properties in Self-Supervised Vision Transformers - Paper Explained
Aleksa Gordić - The AI Epiphany via YouTube
AI, Data Science & Cloud Certificates from Google, IBM & Meta
Future-Proof Your Career: AI Manager Masterclass
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This video explains DINO, a self-supervised learning method for vision transformers, using visual intuition and a pseudocode walkthrough. It covers multi-crop training, the teacher network, attention maps, emergent segmentation masks, feature quality, ablations, and representation collapse.
Syllabus
DINO main ideas, attention maps explained
DINO explained in depth
Pseudocode walk-through
Multi-crop and local-to-global correspondence
More details on the teacher network
Results
Ablations
Collapse analysis
Features visualized and outro
Taught by
Aleksa Gordić - The AI Epiphany