Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

VQ-GAN - Taming Transformers for High-Resolution Image Synthesis - Paper Explained

Aleksa Gordić - The AI Epiphany via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This paper explanation covers VQ-GAN, a model that combines a modified VQ-VAE with a GPT-2 transformer to synthesize high-resolution images. It discusses perceptual and patch-based adversarial losses, transformer conditioning and training, sampling strategies, and comparisons with related models.

Syllabus

Intro
A high-level VQ-GAN overview
Perceptual loss
Patch-based adversarial loss
Sequence prediction via GPT
Generating high-res images
Loss explained in depth
Training the transformer
Conditioning transformer
Comparisons and results
Sampling strategies
Comparisons and results continued
Rejection sampling with ResNet or CLIP
Receptive field effects
Comparisons with DALL-E

Taught by

Aleksa Gordić - The AI Epiphany

Reviews

Start your review of VQ-GAN - Taming Transformers for High-Resolution Image Synthesis - Paper Explained

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.