Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

OpenAI CLIP - Connecting Text and Images - Paper Explained

Aleksa Gordić - The AI Epiphany via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course explains the CLIP paper’s contrastive learning approach for creating transferable image representations from image-text pairs. It examines zero-shot classification, prompt ensembling, embedding quality, robustness to distribution shift, and limitations such as poor MNIST performance.

Syllabus

OpenAI's CLIP
Detailed explanation of the method
Comparision with SimCLR
How does the zero-shot part work
WIT dataset
Why this method, hint efficiency
Zero-shot - generalizing to new tasks
Prompt programming and ensembling
Zero-shot perf
Few-shot comparison with best baselines
How good the zero-shot classifier is?
Compute error correlation
Quality of CLIP's embedding space
Robustness to distribution shift
Limitations MNIST failure
A short recap

Taught by

Aleksa Gordić - The AI Epiphany

Reviews

Start your review of OpenAI CLIP - Connecting Text and Images - Paper Explained

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.