Skip to content
Original
Sebastian Raschka· @rasbt · X·· 3 days agoAI score38

The title and summary in the selected language are awaiting translation.

AI overview

Sebastian Raschka 发布"Reasoning from scratch"系列第 6 期,讲解并实现可验证奖励强化学习(RLVR)与组相对策略优化(GRPO)。

Full text

Reasoning from scratch, round number 6!
An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO).

00:00 Introduction
01:54 What makes a reasoning model different?
04:25 Reasoning traces and model capability
08:29 Accuracy and format rewards
11:34 Aha moments and DeepSeek-R1 training
14:41 Reasoning effort and answer length
18:38 RLHF and RLVR
23:04 GRPO vs. PPO
26:40 GRPO explained with a cooking analogy
31:43 The KL term and simplified GRPO
35:04 Loading the pretrained model
36:07 Loading the MATH training data
39:26 Sampling model responses
46:30 Computing verifiable rewards
49:55 Computing advantages
51:54 Token and sequence log probabilities
55:29 Implementing sequence log probabilities
57:37 Fixing the inference-mode error
1:02:24 Computing the GRPO loss
1:04:37 Putting the GRPO step together
1:09:19 The GRPO training loop
1:12:57 Training settings, logging, and checkpoints
1:17:24 Running training and inspecting outputs
1:19:28 Loading and evaluating checkpoints
1:22:33 MATH-500 results and training stability
1:24:05 Memory requirements and next steps

Source: Sebastian Raschka · x.com