Class Central is learner-supported. When you buy through links on our site, we may earn an affiliate commission.

YouTube

Ultimate Guide to Scaling ML Models - Megatron-LM - ZeRO - DeepSpeed - Mixed Precision

Aleksa Gordić - The AI Epiphany via YouTube

Overview

Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course explains the distributed-training techniques used to scale large machine-learning models, including data, model, pipeline, and tensor parallelism, activation checkpointing, mixed-precision training, and the ZeRO optimizer.

Syllabus

Intro to training Large ML models trillions of params!
sponsored AssemblyAI's speech transcription API
Data parallelism
Pipeline/model parallelism
Megatron-LM paper tensor/model parallelism
Splitting the MLP block vertically
Splitting the attention block vertically
Activation checkpointing
Combining data + model parallelism
Scaling is all you need and 3D parallelism
Mixed precision training paper
Single vs half vs bfloat number formats
Storing master weights in single precision
Loss scaling
Arithmetic precision matters
ZeRO optimizer paper DeepSpeed library
Partitioning is all you need?
Where did all the memory go?
Outro

Taught by

Aleksa Gordić - The AI Epiphany

Reviews

Start your review of Ultimate Guide to Scaling ML Models - Megatron-LM - ZeRO - DeepSpeed - Mixed Precision

Never Stop Learning.

Get personalized course recommendations, track subjects and courses with reminders, and more.

Someone learning on their laptop while sitting on the floor.