AI, Data Science & Cloud Certificates from Google, IBM & Meta
Power BI Fundamentals - Create visualizations and dashboards from scratch
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course walks through the metaseq codebase behind Meta’s OPT-175B large language model, covering environment setup, transformer construction, dummy data, and the training loop. It also explains mixed-precision training, loss scaling, gradient handling, and CUDA/C++ kernels.
Syllabus
Intro - open pretrained transformer
Setup creating the cond env
Setup patch the code
Collecting train script arguments
Training script walk-through
Constructing a dummy task
Building the transformer model
CUDA kernels C++ code
Preparing a dummy dataset
Training loop
Zero grad loss scaling
Forward pass through a transformer
IMPORTANT loss, scaling, mixed precision, error handling
Outro
Taught by
Aleksa Gordić - The AI Epiphany