Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

AudioGen- Textually Guided Audio Generation - Paper Explained

Aleksa Gordić - The AI Epiphany via YouTube

Overview

Coursera Flash Sale
40% Off Coursera Plus for 3 Months!
Grab it
Dive deep into the world of text-guided audio synthesis with this comprehensive video explanation of the "AudioGen: Textually Guided Audio Generation" paper. Explore the challenges of text-to-audio conversion, compare AudioGen with VQ-GAN and SoundStream, and gain insights into audio representation, LSTM networks, and complex-valued STFTs. Learn about audio language modeling, multi-stream audio inputs, data augmentation techniques, and examine the impressive results of this innovative approach to audio generation.

Syllabus

Intro
Why is text-to-audio hard?
Comparison with VQ-GAN
Comparison with SoundStream
AudioGen overview
Deep dive: audio representation, LSTM
Losses explained
Complex-valued STFTs
Audio Language Modeling
Multi-stream audio inputs
Data and augmentations
Results
Outro

Taught by

Aleksa Gordić - The AI Epiphany

Reviews

Start your review of AudioGen- Textually Guided Audio Generation - Paper Explained

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.