Перейти к содержимому
Оригинал
Sebastian Raschka· @rasbt · X·· 3 дня назадОценка ИИ38

Заголовок и краткое изложение на выбранном языке ожидают перевода.

Краткий обзор ИИ

Sebastian Raschka 发布"Reasoning from scratch"系列第 6 期,讲解并实现可验证奖励强化学习(RLVR)与组相对策略优化(GRPO)。

Полный текст

Полный текст на выбранном языке ожидает перевода. Пока показан оригинал.

Reasoning from scratch, round number 6!
An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO).

00:00 Introduction
01:54 What makes a reasoning model different?
04:25 Reasoning traces and model capability
08:29 Accuracy and format rewards
11:34 Aha moments and DeepSeek-R1 training
14:41 Reasoning effort and answer length
18:38 RLHF and RLVR
23:04 GRPO vs. PPO
26:40 GRPO explained with a cooking analogy
31:43 The KL term and simplified GRPO
35:04 Loading the pretrained model
36:07 Loading the MATH training data
39:26 Sampling model responses
46:30 Computing verifiable rewards
49:55 Computing advantages
51:54 Token and sequence log probabilities
55:29 Implementing sequence log probabilities
57:37 Fixing the inference-mode error
1:02:24 Computing the GRPO loss
1:04:37 Putting the GRPO step together
1:09:19 The GRPO training loop
1:12:57 Training settings, logging, and checkpoints
1:17:24 Running training and inspecting outputs
1:19:28 Loading and evaluating checkpoints
1:22:33 MATH-500 results and training stability
1:24:05 Memory requirements and next steps

Источник: Sebastian Raschka · x.com