Learn AI, Data Science & Business — Earn Certificates That Get You Hired
The Fastest Way to Become a Backend Developer Online
Overview
Google, IBM & Meta Certificates — All 10,000+ Courses at 40% Off
One annual plan covers every course and certificate on Coursera. 40% off for a limited time.
Get Full Access
Explore groundbreaking research revealing how a small fraction of "high-entropy" tokens serve as critical decision points in reinforcement learning for large language model reasoning. Discover how focusing learning updates exclusively on these minority forking tokens enables LLMs to achieve comparable or superior reasoning performance compared to training on all tokens, particularly at scale. Learn about the findings from "Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning" by researchers from Qwen Team at Alibaba Inc. and LeapLab at Tsinghua University, including insights into DAPO (Direct Advantage Policy Optimization) reinforcement learning techniques and their implications for AI reasoning capabilities.
Syllabus
High-Entropy Tokens: 20% Control AI Reasoning (DAPO RL)
Taught by
Discover AI