The Most Addictive Python and SQL Courses
Learn AI, Data Science & Business — Earn Certificates That Get You Hired
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This visual course explains the Transformer architecture for language modeling, covering feed-forward networks, self-attention, tokenization, embeddings, output probabilities, and next-token generation.
Syllabus
Intro
The Architecture of the Transformer
Model Training
Transformer LM Component 1: FFNN
Transformer LM Component 2: Self-Attention
Tokenization: Words to Token Ids
Embedding: Breathe meaning into tokens
Projecting the Output: Turning Computation into Language
Final Note: Visualizing Probabilities
Taught by
Jay Alammar