Get 20% off all career paths from fullstack to AI
Start speaking a new language. It’s just 3 weeks away.
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This course implements DeepSeek V3 from scratch in Python. It covers attention mechanisms, KV caching, Multihead Latent Attention, RoPE, mixture-of-experts gating, and transformer blocks.
Syllabus
⌨️ 0:00:00 Intro
⌨️ 0:01:40 Attention Mechanism
⌨️ 0:13:34 Query, Key, Value
⌨️ 0:34:11 KV Cache
⌨️ 0:39:06 Multihead Latent Attention MLA
⌨️ 0:58:53 Coding MLA
⌨️ 1:28:41 RoPE
⌨️ 1:55:44 Coding KV Cache
⌨️ 2:00:25 MLA forward
⌨️ 2:28:24 MoE, Gate
⌨️ 2:49:25 Gate code
⌨️ 3:09:10 MoE code
⌨️ 3:28:36 Transformer Blocks
Taught by
freeCodeCamp.org