Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

DALL-E - Zero-Shot Text-to-Image Generation - Paper Explained

Aleksa Gordić - The AI Epiphany via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course explains the DALL-E research paper and its two-stage approach: VQ-VAE image compression followed by an autoregressive transformer over text and image tokens. It also discusses training details, evaluation, filtering with CLIP, and emergent image-to-image translation.

Syllabus

What is DALL-E?
VQ-VAE blur problems
transformers, transformers, transformers!
Stage 1 and Stage 2 explained
Stage 1 VQ-VAE recap
Stage 2 autoregressive transformer
Some notes on ELBO
VQ-VAE modifications
Stage 2 in-depth
Results
Engineering, engineering, engineering
Automatic filtering via CLIP
More results
Additional image to image translation examples

Taught by

Aleksa Gordić - The AI Epiphany

Reviews

Start your review of DALL-E - Zero-Shot Text-to-Image Generation - Paper Explained

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.