Skip to content

#Code review

0 items today
10/6Tue
  1. Tproger · Программирование58

    GitHub 与 Microsoft 开放 ReviewBench 评测 AI 代码评审

    GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。

    Awaiting translation

  2. Habr · Вайбкодинг62

    一位 Rust 开发者为什么仍然害怕用 AI 写代码

    Rust 开发者 NikTimf 在 Habr 撰文说,自己仍然害怕用 AI 写代码,原因不是生成质量差,而是生成量太大、后续没人真正看懂。他列举了具体代价:多个智能体之间要反复传递上下文,答案冲突时还得自己判断谁对;公司只允许本地或自研模型时,用惯强模型的人很难退回;同事充当 meat proxy 转发 AI 答案,理解任务和推进实现的活仍落在自己身上。

    Awaiting translation

  3. GitHub Blog · Copilot66

    GitHub 发布 AI 代码评审开放基准 ReviewBench

    GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。

    Awaiting translation

    Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。

10/5Mon
10/4Sun
  1. Ben Holmes58

    Ben Holmes 公开了驱动其软件工厂的多智能体系统:需求来自 Linear issue 或 Slack 讨论,分诊智能体先调研并决定是直接实现还是提问,实现智能体负责构建、必要时先与子智能体协作产出 spec,验证子智能体对实现结果做端到端测试,代码评审智能体与实现智能体循环几轮后再交人工评审上线,监控自动化则响应告警并创建 issue。

    Awaiting translation

10/1Thu
  1. 冰河技术20

    领导说三个月上微服务,发掉了半斤,进度还在原地转圈

    一位开发者复盘自己按功能模块拆微服务失败的经历:一个下单流程要调八个服务,改订单状态会连带库存、支付、物流报错,凌晨被报警叫醒成常态。同事用领域驱动设计的“拆服务四步法”重新拆分后,下单响应时间从平均1.2秒降到350毫秒。文章给出事件风暴、聚合根设计、按限界上下文拆服务等步骤,并附 Java 代码示例。

    Awaiting translation

9/28Mon
9/26Sat
9/24Thu
9/23Wed
9/15Tue
9/14Mon
  1. Addy Osmani · Blog80

    Addy Osmani on engineering methods for bringing AI agents into legacy codebases

    Addy Osmani suggests that when introducing AI agents into brownfield codebases, you should first make hidden constraints visible and make cheap changes trustworthy. He recommends dividing code into green, yellow, and red zones: green zones with solid tests let agents move in small, fast steps; yellow zones require writing characterization tests first; and red zones involving sensitive logic like authentication, billing, and permissions must have humans involved step by step. The zones are drawn by hand, and a yellow zone can only be upgraded to green once characterization tests exist and the module owner has reviewed the first batch of changes.

    Why it matters: The author turns the constraints of bringing agents into an old codebase into actionable rules—zoning, characterization tests, and migration units—and cites migration data from several companies as reference.

9/12Sat
  1. Simon Willison · Coding Agents0

    Boris Cherny:Claude 写的生产代码应有比人类更高的门槛

    Anthropic 的 Boris Cherny 表示,Claude 编写的生产代码应比人类编写的代码有更高门槛。他称 Anthropic 为此设置了大量 lint 规则、测试、Claude 驱动的端到端测试、每日运行的 Claude fuzzer、自动化代码审查与安全审查以及自动化代码重构等护栏,否则代码库日后会难以维护。

    Awaiting translation

9/11Fri
9/6Sun
  1. Thorsten Ball · Register Spill15

    Thorsten Ball 谈 AI 写代码:是代码真的差,还是评审者的“天真干预”偏见

    Thorsten Ball 在 Register Spill 的 Joy & Curiosity #98 中借用《反脆弱》里的扁桃体切除研究,提出工程师对 Sol、Fable、Astra 等模型输出的“代码差、注释蠢”评价,可能源于“天真干预主义”偏见——AI 已在 20 分钟内端到端完成前后端改动、内外部文档和测试,并在无头浏览器中跑完全流程、附上录屏为证。

    Awaiting translation

9/5Sat
  1. Vibe Code Textbook · Articles87

    审查 coding agent 的 diff:检查清单最先抓到什么

    作者给出一份按危害排序的十项 agent diff 检查清单,依次看被删除的测试、被跳过或弱化的断言、宽泛异常捕获、新增依赖、任务范围外文件、CI 配置改动、疑似密钥、遗留标记和净删除超过 40 行的文件。

    Awaiting translation

    Why it matters: 给出按危害排序的十项 agent diff 检查清单,并附可复用的扫描脚本与行号定位。

9/2Wed
  1. Hacker News · Agent Skills78

    mattpocock releases AI coding Agent Skills built for real engineering

    Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.

    Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.

8/28Fri
8/18Tue
8/13Thu
  1. Augment Code · Blog62

    Augment Code Expands Cosmos: Turning Code Review into an Agentic PR-to-Merge Loop

    Augment Code has extended its Cosmos review system from code review to a full PR-to-merge loop, adding four capabilities: Verifier, PR Fixer, Review Dashboard, and cosmos approve. Dedicated Experts handle risk analysis, line-by-line correctness review, design review, runtime verification, and fixes.

    Why it matters: Augment has expanded code review into a PR-to-merge loop covering fixes, verification, and approval, giving readers a way to judge how multi-agent division of labor plays out in practice.

8/8Sat
  1. Johnny Butler · Agentic Engineering64

    过早的错误处理暴露了 Agent 把功能建错了顺序

    作者在评审数千个编码 Agent 的改动后发现,过早加入 rescue 块、兜底逻辑和日志,往往说明功能构建顺序错了。Agent 倾向横向铺开数据库、服务、API、校验和 UI,提前设想完整系统,这与 SWE-bench 只衡量补丁能否解决限定问题并通过测试的评估方式相符。

    Awaiting translation

8/5Wed
  1. Kondasamy Jayaraman · Engineering Blog58

    为什么 JSONL 更适合 AI 智能体工作负载

    作者结合 Cloudflare 编排数千个合并请求代码评审的做法,说明智能体进程为何普遍在 stdout 输出 JSONL。普通 JSON 必须等到闭合括号才能解析,进程崩溃时整份输出作废;JSONL 每行是独立合法对象,崩溃后已写出的行仍可解析,也便于追加、流式读取和按行拆分给多个 worker。

    Awaiting translation

8/4Tue
8/2Sun
7/20Mon
  1. OpenAI Developer Blog · Codex71

    Codex Code Review now supports custom review rules in AGENTS.md

    OpenAI has added custom repository rules to Codex Code Review: you can put review guidelines in AGENTS.md, and Codex applies them during review and cites where each one came from in its findings. In OpenAI's own evaluation, the rule-guided version caught 98% of the required custom issues, versus 58.3% for the baseline. The guidance is to start with non-obvious invariants like compatibility requirements and data boundaries, put repo-level rules in the root directory and service-level rules in the corresponding directory, and leave formatting and mechanical checks to CI.

    Why it matters: OpenAI lays out the capabilities, the syntax, and the evaluation data for Codex Code Review custom rules, so you can judge how to bake your team's review experience into AGENTS.md.

7/18Sat
  1. Thorsten Ball · Register Spill22

    Joy & Curiosity #92:Amp 推出订阅与智能体间通信

    Amp 现已推出订阅服务,可与 ChatGPT 订阅搭配获得无限 GPT-5.6 token;同时 Amp 上线智能体间通信,智能体能在任意 Amp 实例或 orb 中派生其他智能体并互发消息与文件。作者还分享了新一季 Raising An Agent 播客、与 Evan Phoenix 等人的对谈,以及 antirez 关于“控制想法而非代码”的观点。

    Awaiting translation

7/15Wed
  1. Johnny Butler · Agentic Engineering60

    Agent 写出的 Pull Request 比人类更好

    作者评审了一个 Agent 提交的 Pull Request,认为它又快又好,原因不在模型有多聪明,而在于质量门槛被写进了循环内部。这个 PR 在请求人工介入前就说明了改动了什么、刻意没动什么、该重点审查哪里、遵循了哪些规范,并附上了已运行的验证证据。作者由此提出,给 Agent 一个明确的完成定义和必须自证的标准,它就会为通过标准而优化,速度不是靠降低门槛换来的。

    Awaiting translation

7/8Wed
  1. AI Hero · Skills Updates65

    AI Hero skills repo ships v1.1: adds /wayfinder, renames /to-spec and /to-tickets

    AI Hero's skills repo ships v1.1, renaming /to-prd to /to-spec, merging /to-plan and /to-issues into /to-tickets, and adding new Skills like /wayfinder, /research, and /prototype.

    Why it matters: The author walks through the full Skill flow from grilling to deployment and gives the migration commands for the renames, the merge, and the new /wayfinder—useful for anyone building an AI development workflow.

7/3Fri
  1. Johnny Butler · Agentic Engineering31

    结对编程能否解决 AI 代码 PR 审查瓶颈?

    结对编程被一些团队用来把同行评审嵌入开发过程,让评审在工作进行时完成,从而缓解 AI 辅助交付带来的 PR 审查瓶颈。其机制在于评审信心产生于编码过程中,而非事后检查;AI 智能体产出代码更快更多,若判断全部堆到 PR 阶段,瓶颈只会加剧。应对之策是把标准、质疑和判断前移到工作本身,让证据随 PR 一起到达。

    Awaiting translation

  2. Lovable · Blog88

    花掉 8.5 万美元 token 后,我在 Lovable 扩展智能体编程的经验

    Lovable 一名工程师从今年 1 月到 6 月把个人 token 花费从每月约 600 美元推到 5 月的约 2.5 万美元、累计约 8.5 万美元,同时把每周合并 PR 数从 20-30 个提升到 150 个以上。

    Awaiting translation

    Why it matters: 作者公开了自己每月约 2.5 万美元 token 的智能体开发配置,包括风险分级、多智能体评审和上下文管理,可迁移到其他团队。

6/23Tue
  1. Johnny Butler · Agentic Engineering38

    你的 AI 代码审查工具来得太晚了

    让 AI 智能体先写代码、再用 CodeRabbit、Greptile 这类审查工具事后清理,可能已经太晚:等审查工具看到 PR 时,智能体早已定下实现形态、假设和测试策略。昂贵的错误通常发生在 PR 之前,比如误读系统、切片过大或沿用错误模式,事后打磨无法纠正方向。审查工具应作为检查环节而非交付模式,把结果、约束、验收标准和验证前置到工作流程中。

    Awaiting translation

6/19Fri
  1. Hacker News · AI Code Review 讨论40

    Ask HN:大家都在用哪些 AI 辅助代码审查工具?

    一个约 40 人规模的开发团队正在评估 AI 辅助代码审查工具,在开始一系列免费试用前向社区征集经验。提问者想了解大家使用哪些工具或服务、是否只用于代码审查,还是也用于事件响应、分支管理等场景,以及选择原因和优缺点。

    Awaiting translation

6/17Wed
6/9Tue
  1. Johnny Butler · Agentic Engineering34

    Governed PRs:如何在智能体速度下守住质量门槛

    Governed PRs 通过让 AI 智能体在 PR 中展示其应用的 Playbooks Applied 部分,公开所参照的标准、遵循的仓库规则以及仍需人工判断之处,从而在智能体高速产出 PR 的同时守住质量门槛。作者称在自己的 SDF 实践中,这一做法显著缩短了交付周期且未降低质量标准,因为评审从上下文而非考古式追溯开始。

    Awaiting translation