Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

Do Vision Transformers See Like Convolutional Neural Networks - Paper Explained

Aleksa Gordić - The AI Epiphany via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course explains a research paper comparing vision transformers and convolutional neural networks. It examines receptive fields, feature evolution, skip connections, spatial information, token flow, and the effect of training data using visual and geometric intuition.

Syllabus

Intro
Contrasting features in ViTs vs CNNs
Global vs Local receptive fields
Data matters, mr. obvious
Contrasting receptive fields
Data flow through CLS vs spatial tokens
Skip connections matter a lot in ViTs
Spatial information is preserved in ViTs
Features evolution with the amount of data
Outro

Taught by

Aleksa Gordić - The AI Epiphany

Reviews

Start your review of Do Vision Transformers See Like Convolutional Neural Networks - Paper Explained

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.