Python, Prompt Engineering, Data Science — Build the Skills Employers Want Now
Learn Backend Development Part-Time, Online
Overview
Google, IBM & Meta Certificates — All 10,000+ Courses at 40% Off
One annual plan covers every course and certificate on Coursera. 40% off for a limited time.
Get Full Access
Explore a comprehensive video explanation of the Reinforced Self-Training (ReST) method for language modeling. Delve into how ReST utilizes a bootstrap-like approach to generate its own extended dataset, training on increasingly high-quality subsets to enhance its reward system. Understand the efficiency advantages of ReST compared to Online Reinforcement Learning techniques like PPO, including its ability to reuse generated data multiple times. Examine the paper's abstract, which outlines ReST's application in machine translation and its potential to significantly improve translation quality. Learn about the authors behind this innovative approach and their findings on ReST's compute and sample efficiency in improving large language models through alignment with human preferences.
Syllabus
Reinforced Self-Training (ReST) for Language Modeling (Paper Explained)
Taught by
Yannic Kilcher