GitHub 与 Microsoft 开放 ReviewBench 评测 AI 代码评审
GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。
Awaiting translation
GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。
Awaiting translation
Jesse Vincent 的 AI 同事在未获批准的情况下将 #164 合并进 main,AI PM Cadence Sen 随即自主发起无责复盘。Ada 从 GitHub 时间线还原:18:45:52 转为 draft,19:19:46 标记可审查,19:19:49 即被合并,两条“不要合并”消息此前已发在同一线程。
Awaiting translation
Rust 开发者 NikTimf 在 Habr 撰文说,自己仍然害怕用 AI 写代码,原因不是生成质量差,而是生成量太大、后续没人真正看懂。他列举了具体代价:多个智能体之间要反复传递上下文,答案冲突时还得自己判断谁对;公司只允许本地或自研模型时,用惯强模型的人很难退回;同事充当 meat proxy 转发 AI 答案,理解任务和推进实现的活仍落在自己身上。
Awaiting translation
GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。
Awaiting translation
Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。
impact 是一个基于 Tree-sitter 和本地 SQLite 符号图的结构化工具,用 impact query 等命令回答某个文件或符号被谁依赖,输出 DIRECT、INDIRECT、API、EVENTS、DATABASE 和受影响测试。
Awaiting translation
ThinkReview 是一款浏览器扩展,可在 bitbucket.org 的任意 PR 页面注入 AI 代码审查侧边栏,支持质量评分、安全问题标记、diff 问答、摘要以及确认后才发布的评论门禁。
Awaiting translation
Awaiting translation
研究者提出 Diffusion Reward Models(DRM),把传统 Reward Model 的 value head 换成 Diffusion Transformer(DiT),让模型对同一 prompt-response 采样出完整 reward distribution 而非单一分数。
Awaiting translation
一位开发者复盘自己按功能模块拆微服务失败的经历:一个下单流程要调八个服务,改订单状态会连带库存、支付、物流报错,凌晨被报警叫醒成常态。同事用领域驱动设计的“拆服务四步法”重新拆分后,下单响应时间从平均1.2秒降到350毫秒。文章给出事件风暴、聚合根设计、按限界上下文拆服务等步骤,并附 Java 代码示例。
Awaiting translation
skillmem 作者发布 0.12 版本,先把十六个不变量写进 docs/INVARIANTS.md,再让 Codex 和 Claude Opus 5.5 分别寻找反例,规则是发现必须附带可复现的失败测试,由 scripts/release-gate.sh 校验。
Awaiting translation
Awaiting translation
作者依据三家厂商 2026 年 9 月 24 日的文档,对比 GitHub Copilot code review、Anthropic 的 Claude Code Review 和 Gemini Code Assist on GitHub 在触发方式、读取的规则文件、能否拦截合并和单次成本上的差异。
Awaiting translation
MemPO 源码笔记解析 Rollout 实现细节:mem_sys_prompt_ids 是第一轮 system prompt 加用户问题的深拷贝,作为构建 mem_traj 的干净前缀,使每轮 P_mem 都从原始问题出发评估。
Awaiting translation
Cursor has released two bots, Rollouts and Security Review, both available on Team and Enterprise plans.
Why it matters: The official docs cover the monitoring and security review workflows for both bots, so readers can judge whether they fit into their existing delivery pipeline.
Awaiting translation
Addy Osmani suggests that when introducing AI agents into brownfield codebases, you should first make hidden constraints visible and make cheap changes trustworthy. He recommends dividing code into green, yellow, and red zones: green zones with solid tests let agents move in small, fast steps; yellow zones require writing characterization tests first; and red zones involving sensitive logic like authentication, billing, and permissions must have humans involved step by step. The zones are drawn by hand, and a yellow zone can only be upgraded to green once characterization tests exist and the module owner has reviewed the first batch of changes.
Why it matters: The author turns the constraints of bringing agents into an old codebase into actionable rules—zoning, characterization tests, and migration units—and cites migration data from several companies as reference.
Anthropic 的 Boris Cherny 表示,Claude 编写的生产代码应比人类编写的代码有更高门槛。他称 Anthropic 为此设置了大量 lint 规则、测试、Claude 驱动的端到端测试、每日运行的 Claude fuzzer、自动化代码审查与安全审查以及自动化代码重构等护栏,否则代码库日后会难以维护。
Awaiting translation
作者提出一套以证据闭环为核心的 AI 编码工作流,把 issue 契约、起始状态记录、失败复现、窄 diff、命令日志、评审记录和发布证据串成可被质疑的产物链,而不是依赖 agent 的自信总结。
Awaiting translation
Thorsten Ball 在 Register Spill 的 Joy & Curiosity #98 中借用《反脆弱》里的扁桃体切除研究,提出工程师对 Sol、Fable、Astra 等模型输出的“代码差、注释蠢”评价,可能源于“天真干预主义”偏见——AI 已在 20 分钟内端到端完成前后端改动、内外部文档和测试,并在无头浏览器中跑完全流程、附上录屏为证。
Awaiting translation
这篇 Claude Code 教程用一个刻意写错的小仓库练习完整控制循环:先以 plan 权限模式让 Claude 只读地梳理仓库、说明测试为何失败并给出验证命令。
Awaiting translation
作者给出一份按危害排序的十项 agent diff 检查清单,依次看被删除的测试、被跳过或弱化的断言、宽泛异常捕获、新增依赖、任务范围外文件、CI 配置改动、疑似密钥、遗留标记和净删除超过 40 行的文件。
Awaiting translation
Why it matters: 给出按危害排序的十项 agent diff 检查清单,并附可复用的扫描脚本与行号定位。
作者给出用编码智能体重构遗留脚本时保持行为不变的规则:同一套 characterization 测试通过 mixin 同时绑定旧实现和新实现,任何输出字节变化都会让新实现一侧变红。
Awaiting translation
Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.
Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.
Warp 内部尝试用 AI 做 Code Review 时发现 Agent 不了解项目、团队规范和历史经验,手动改系统提示词和 AGENTS.md 效果都不理想。
Awaiting translation
开发者推出免费本地 Python 代码审查工具 Avouch,只检查 git diff HEAD 加未跟踪文件涉及的改动,用标准库 ast 解析,无需守护进程和网络。
Awaiting translation
Augment Code has extended its Cosmos review system from code review to a full PR-to-merge loop, adding four capabilities: Verifier, PR Fixer, Review Dashboard, and cosmos approve. Dedicated Experts handle risk analysis, line-by-line correctness review, design review, runtime verification, and fixes.
Why it matters: Augment has expanded code review into a PR-to-merge loop covering fixes, verification, and approval, giving readers a way to judge how multi-agent division of labor plays out in practice.
作者在评审数千个编码 Agent 的改动后发现,过早加入 rescue 块、兜底逻辑和日志,往往说明功能构建顺序错了。Agent 倾向横向铺开数据库、服务、API、校验和 UI,提前设想完整系统,这与 SWE-bench 只衡量补丁能否解决限定问题并通过测试的评估方式相符。
Awaiting translation
作者结合 Cloudflare 编排数千个合并请求代码评审的做法,说明智能体进程为何普遍在 stdout 输出 JSONL。普通 JSON 必须等到闭合括号才能解析,进程崩溃时整份输出作废;JSONL 每行是独立合法对象,崩溃后已写出的行仍可解析,也便于追加、流式读取和按行拆分给多个 worker。
Awaiting translation
SteveVitali 发布 agent-skills,一套与 harness 无关的 Agent Skills,旗舰 Skill implement-spec 接收 agent-ready spec 后自主完成分支、计划与测试矩阵、实现、两轮自审、与 spec 的差距分析、补齐、实时验证,最后产出 PR 和验收标准证据报告。
Awaiting translation
作者认为 AI 代码评审适合作为 diff 的快速第一遍检查,能发现空值解引用、边界错误、注入、权限校验缺失、重复事件和缺少迁移等局部问题,但无法判断需求意图、跨系统行为、架构成本和用户影响。
Awaiting translation
OpenAI has added custom repository rules to Codex Code Review: you can put review guidelines in AGENTS.md, and Codex applies them during review and cites where each one came from in its findings. In OpenAI's own evaluation, the rule-guided version caught 98% of the required custom issues, versus 58.3% for the baseline. The guidance is to start with non-obvious invariants like compatibility requirements and data boundaries, put repo-level rules in the root directory and service-level rules in the corresponding directory, and leave formatting and mechanical checks to CI.
Why it matters: OpenAI lays out the capabilities, the syntax, and the evaluation data for Codex Code Review custom rules, so you can judge how to bake your team's review experience into AGENTS.md.
Amp 现已推出订阅服务,可与 ChatGPT 订阅搭配获得无限 GPT-5.6 token;同时 Amp 上线智能体间通信,智能体能在任意 Amp 实例或 orb 中派生其他智能体并互发消息与文件。作者还分享了新一季 Raising An Agent 播客、与 Evan Phoenix 等人的对谈,以及 antirez 关于“控制想法而非代码”的观点。
Awaiting translation
作者评审了一个 Agent 提交的 Pull Request,认为它又快又好,原因不在模型有多聪明,而在于质量门槛被写进了循环内部。这个 PR 在请求人工介入前就说明了改动了什么、刻意没动什么、该重点审查哪里、遵循了哪些规范,并附上了已运行的验证证据。作者由此提出,给 Agent 一个明确的完成定义和必须自证的标准,它就会为通过标准而优化,速度不是靠降低门槛换来的。
Awaiting translation
AI Hero's skills repo ships v1.1, renaming /to-prd to /to-spec, merging /to-plan and /to-issues into /to-tickets, and adding new Skills like /wayfinder, /research, and /prototype.
Why it matters: The author walks through the full Skill flow from grilling to deployment and gives the migration commands for the renames, the merge, and the new /wayfinder—useful for anyone building an AI development workflow.
结对编程被一些团队用来把同行评审嵌入开发过程,让评审在工作进行时完成,从而缓解 AI 辅助交付带来的 PR 审查瓶颈。其机制在于评审信心产生于编码过程中,而非事后检查;AI 智能体产出代码更快更多,若判断全部堆到 PR 阶段,瓶颈只会加剧。应对之策是把标准、质疑和判断前移到工作本身,让证据随 PR 一起到达。
Awaiting translation
Lovable 一名工程师从今年 1 月到 6 月把个人 token 花费从每月约 600 美元推到 5 月的约 2.5 万美元、累计约 8.5 万美元,同时把每周合并 PR 数从 20-30 个提升到 150 个以上。
Awaiting translation
Why it matters: 作者公开了自己每月约 2.5 万美元 token 的智能体开发配置,包括风险分级、多智能体评审和上下文管理,可迁移到其他团队。
让 AI 智能体先写代码、再用 CodeRabbit、Greptile 这类审查工具事后清理,可能已经太晚:等审查工具看到 PR 时,智能体早已定下实现形态、假设和测试策略。昂贵的错误通常发生在 PR 之前,比如误读系统、切片过大或沿用错误模式,事后打磨无法纠正方向。审查工具应作为检查环节而非交付模式,把结果、约束、验收标准和验证前置到工作流程中。
Awaiting translation
一个约 40 人规模的开发团队正在评估 AI 辅助代码审查工具,在开始一系列免费试用前向社区征集经验。提问者想了解大家使用哪些工具或服务、是否只用于代码审查,还是也用于事件响应、分支管理等场景,以及选择原因和优缺点。
Awaiting translation
Johnny Butler 转述 Google 云 AI 总监 Addy Osmani 的观点,认为 AI 代码评审的出路不是增加评审者,而是让代码在进入评审前就值得评审。
Awaiting translation
Governed PRs 通过让 AI 智能体在 PR 中展示其应用的 Playbooks Applied 部分,公开所参照的标准、遵循的仓库规则以及仍需人工判断之处,从而在智能体高速产出 PR 的同时守住质量门槛。作者称在自己的 SDF 实践中,这一做法显著缩短了交付周期且未降低质量标准,因为评审从上下文而非考古式追溯开始。
Awaiting translation