跳到正文
原文
Sebastian Raschka· @rasbt · X·· 10天前AI 评分22
AI 导读

Sebastian Raschka 发布《Reasoning from scratch》第 5 期,讲解对数概率评分与自我精炼。内容涵盖加载预训练 LLM、用 PyTorch 计算 token 概率与序列似然、构建基于规则的评分器,以及自我精炼循环的实现,并在 MATH-500 上给出评测结果。

正文

当前语言的正文正在等待翻译,暂时显示原文。

Reasoning from scratch, round number 5! This time, talking about log-probability scoring (also a great fundamental concept for loss functions like cross-entropy in pre-training and distillation) and self-refinement.

00:00 Introduction and inference-time scaling recap
05:02 Loading the pretrained LLM
08:00 Comparing and scoring model answers
10:18 Building a rule-based scorer
17:53 Token probabilities and sequence likelihood
26:47 Computing token probabilities in PyTorch
30:12 Token indexing and shifted targets
37:27 Log probabilities and numerical stability
45:57 Scoring answers with average log probabilities
56:24 How self-refinement works
59:07 Generating critiques and revised answers
1:01:00 Implementing the self-refinement loop
1:05:57 MATH-500 evaluation results
1:07:35 Takeaways and next steps

来源:Sebastian Raschka · x.com