Transformers that Transform Well Enough to Support Near-Shallow Architectures - Stanford CS25 Lecture
Stanford University via YouTube
Free courses from frontend to fullstack and AI
AI, Data Science & Cloud Certificates from Google, IBM & Meta
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
This talk introduces precision language models, which use specialized self-attention architectures and non-random parameter initialization to train language models efficiently with limited resources. It also describes a developing application that runs these models on microprocessors as controllers for small electronic devices.
Syllabus
Stanford CS25: V4 I Transformers that Transform Well Enough to Support Near-Shallow Architectures
Taught by
Stanford Online