本课程面向具备机器学习基础的学生,系统讲授从经典策略梯度到大语言模型后训练的策略优化方法。学生将理解优势估计、近端更新和组相对学习信号的核心思想,能够比较REINFORCE、TRPO、PPO、DPO、GRPO及其改进方法,识别训练稳定性问题,并建立阅读前沿论文和分析实际训练流程的能力
Learn Excel and Financial Modeling the Way Finance Teams Actually Use Them
AI, Data Science & Cloud Certificates from Google, IBM & Meta
Overview
Google, IBM & Meta Certificates – 40% Off
One Coursera Plus subscription covers most Professional Certificates on Coursera.
Unlock All Certificates
Syllabus
- 第一章 REINFORCE与Actor-Critic
- 第二章 TRPO、PPO与DPO
- 第三章 GRPO与DAPO
- 第四章 HA-DW与OPD
Taught by
byk and wrj