Boris Cherny 用 Opus 5.5 配合 Lean 对 Claude Agent SDK 做形式化验证,几段简短提示词换来 16 个 PR,修复了多个 bug 和竞态条件。
Awaiting translation
Boris Cherny 用 Opus 5.5 配合 Lean 对 Claude Agent SDK 做形式化验证,几段简短提示词换来 16 个 PR,修复了多个 bug 和竞态条件。
Awaiting translation
Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.
Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.
Paper Compute 工程师用 TypeSafe 的 Jev 做智能体会话标签实验,在 1,781 条工程师对话轮次上,Jev 1.13.0 中位响应 188 ms,比 GPT-4o mini 的 1,023 ms 和 Claude Haiku 4.5 的 1,050 ms 快约 5.4 倍和 5.6 倍,本地 Qwen3 8B 为 2,311 ms。
Awaiting translation
Foremerge 是一个位于 git 之上的本地协调层,不改变 git 和 worktree 的工作方式。每个智能体在编辑前先发布意图和将要改动的范围及操作,第二个智能体发布时若与已有意图冲突会返回 HIGH 级别的 destructive_vs_additive 提示,检测是确定性的,不依赖评审模型读代码。
Awaiting translation
文章用概率模型算出,失败率 1% 的测试要跑 299 次才有 95% 概率被发现,而单次重跑有 99% 概率变绿,因此重跑通过几乎不提供信息。作者建议智能体循环中不要用重跑覆盖原始失败记录,并把已知不稳定测试单独标记,例如用 pytest-rerunfailures 的 @pytest.mark.flaky(reruns=n) 按测试标注,而不是全局 --reruns 3。
Awaiting translation
作者提出 AI 生成代码的测试应从任务契约出发,攻击生成代码容易忽略的边界,包括无效输入、缺失身份、错误权限、空值与最大值、重复投递、重试、并发请求、schema 兼容、部分失败和回滚,并给出一张覆盖契约、输入、身份、权限、持久化、重复、重试、并发、迁移、外部服务、可观测性和回滚的失败矩阵。
Awaiting translation
Addy Osmani suggests that when introducing AI agents into brownfield codebases, you should first make hidden constraints visible and make cheap changes trustworthy. He recommends dividing code into green, yellow, and red zones: green zones with solid tests let agents move in small, fast steps; yellow zones require writing characterization tests first; and red zones involving sensitive logic like authentication, billing, and permissions must have humans involved step by step. The zones are drawn by hand, and a yellow zone can only be upgraded to green once characterization tests exist and the module owner has reviewed the first batch of changes.
Why it matters: The author turns the constraints of bringing agents into an old codebase into actionable rules—zoning, characterization tests, and migration units—and cites migration data from several companies as reference.
作者用 Astra、Fable 和 Opus 连续数天规划编码一个数万行代码的智能体项目,在 e2e 测试中发现 MCP 工具的响应被重复输出(结构化加普通文本),这个缺陷一直漏到最终验证阶段。
Awaiting translation
Anthropic 的 Boris Cherny 表示,Claude 编写的生产代码应比人类编写的代码有更高门槛。他称 Anthropic 为此设置了大量 lint 规则、测试、Claude 驱动的端到端测试、每日运行的 Claude fuzzer、自动化代码审查与安全审查以及自动化代码重构等护栏,否则代码库日后会难以维护。
Awaiting translation
作者提出一套以证据闭环为核心的 AI 编码工作流,把 issue 契约、起始状态记录、失败复现、窄 diff、命令日志、评审记录和发布证据串成可被质疑的产物链,而不是依赖 agent 的自信总结。
Awaiting translation
Cursor engineer Lauren Tan shares how she gets AI agents to submit and merge PRs on their own: the key is validation—letting the agent run code, capture CPU traces, and open an iOS simulator to check its own work.
Why it matters: Cursor engineers break trust in AI agents down into reusable validation skills, feature maps, and evals, so readers can build their own automated validation workflows.
作者 Johnny Butler 认为,如果编码智能体直到 CI 阶段才发现自己违反了工程规范,反馈就来得太晚,应在智能体实现过程中就运行检查。他建议把测试、linter、类型检查、安全扫描以及团队自定义的依赖规则和复杂度阈值检查交给智能体,让它改完小改动后运行相关检查、修复失败项再重跑,并把结果随代码一起返回,使开发者能用同样的检查复现。
Awaiting translation
这篇 Claude Code 教程用一个刻意写错的小仓库练习完整控制循环:先以 plan 权限模式让 Claude 只读地梳理仓库、说明测试为何失败并给出验证命令。
Awaiting translation
作者提出用六段式 spec 模板(Goal、Non-goals、Interfaces、Files、Verification、Budget)向编码智能体描述任务,并配了一个在交给智能体前检查 spec 的 linter。
Awaiting translation
Why it matters: 给出可直接套用的六段式 spec 模板、tally 实例和配套 linter,读者能据此改造自己交给编码智能体的任务描述。
作者以 tally 这个 CSV 列统计 CLI 为例,给出从一段式 spec 到可运行工具的完整方法:spec 包含目标、非目标、接口、文件、验证和预算六部分,提示词按先列边界情况、再写九条测试、然后实现、跑三条手工检查、最后按 diff 清单评审的顺序发送。
Awaiting translation
SWE-bench 的百分比是智能体编程领域被引用最多的数字,但它比通常被引用的方式要窄得多。文章拆解了一次评测如何产出这个数字:提交是 JSONL 格式的补丁,在 Docker 容器中应用到 base_commit 并运行仓库测试,FAIL_TO_PASS 全部通过且 PASS_TO_PASS 不回归才算 resolved,默认是 pass@1。
Awaiting translation
作者给出一份按危害排序的十项 agent diff 检查清单,依次看被删除的测试、被跳过或弱化的断言、宽泛异常捕获、新增依赖、任务范围外文件、CI 配置改动、疑似密钥、遗留标记和净删除超过 40 行的文件。
Awaiting translation
Why it matters: 给出按危害排序的十项 agent diff 检查清单,并附可复用的扫描脚本与行号定位。
作者给出用编码智能体重构遗留脚本时保持行为不变的规则:同一套 characterization 测试通过 mixin 同时绑定旧实现和新实现,任何输出字节变化都会让新实现一侧变红。
Awaiting translation
文章梳理了 Claude Code、mini-swe-agent、Omnigent 等 harness 提供的硬性停止条件,包括步数、花费、墙钟时间、连续格式错误和显式退出,并建议优先设置美元上限,因为它是唯一随模型规模变化的限制。
Awaiting translation
文章主张让 AI 智能体在写实现前先展示一次失败的测试,因为一次都没红过的测试无法证明它能检测任何行为。作者引用 AgentLens 论文(arXiv:2605.12925)对 2,614 条 OpenHands 轨迹的评测,发现通过轨迹中有 10.7% 属于 Lucky Pass,各模型的幸运通过率在 0.5% 到 23.2% 之间。
Awaiting translation
作者给出用编码智能体给无测试遗留脚本补特征测试(characterization tests)的五步方法:先在 plan 模式做只读调研,找出模块级状态、print 位置和疑似 bug。
Awaiting translation
The author built the space exploration game Void Explorer in Codex with Astra, featuring 2,048 star systems and over 10,000 procedurally generated planets, and shared the full workflow from prompts to architecture, testing, and performance measurement.
Why it matters: Using Astra in Codex, the author built an entire game and showed a transferable collaborative workflow that spans prompts, testing, and performance measurement.
Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.
Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.
作者设计了一套可复现的对比方法,让 Codex 和 Claude Code 在同一个仓库、同一任务契约、同一权限和停止条件下完成五项任务:仓库发现、编辑纪律、测试恢复、评审证据、远程或无人值守执行。
Awaiting translation
这个插件把 Birgitta Boeckeler 在 martinfowler.com 提出的 harness engineering 落地为可安装工具,用确定性工具、智能体评审和周期性熵检查约束 AI 生成的代码。
Awaiting translation
博主用 Qwen3-30B-A3B-Instruct-2507 在租用的 H200 上生成 30 条回答,其中部分用 SynthID-Text 加水印,让读者分辨哪条被水印。首轮 278 人平均得分 3.92/10,重排题目后 73 人平均 3.4/10,接近纯随机猜测的 3.33,说明读者无法识别水印。目前测验累计约 4700 份回答,均值 3.44/10。
Awaiting translation
OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.
Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.
Simon Willison 用 Claude Code for web 做了一个约 150 行、零依赖的 TypeScript 服务原型,基于 Bun 1.4 实验性的 Bun.WebView 提供 shot-scraper 风格的 JSON API,支持执行 JavaScript 和输出 PNG/JPEG/WebP 截图,无需 Puppeteer 或 Playwright。
Awaiting translation
作者提出在 Cursor 中开多个 Agent 窗口只是并行对话,真正要解决的是分工与可检查产物:先用验收卡把需求写成含用户结果、范围、验收标准、所需证据和风险边界的契约,再按 PM、DEV、OPS、QA、EVAL 和人工发布决策划分职责,各自不得替代他人判断。
Awaiting translation
Lovable 用六个月把月访问 4200 万的 lovable.dev 从 Next.js 迁到自家 TanStack Start 托管栈,迁移期间两套框架并行运行,由代理 worker 按路由和用户分流,最终 Next.js 专属代码只占 3%。
Awaiting translation
Why it matters: Lovable 官方复盘把 90 万行代码从 Next.js 迁到自家 TanStack Start 栈,双框架并行、AI 智能体批量迁移和一次 OOM 事故的细节都可直接借鉴。
Wiz 红队智能体利用 Snowflake 开源仓库 GitHub Actions 工作流中的命令注入漏洞,从 CI/CD 环境窃取 Jira API 凭证,这些凭证可读取 Snowflake 工程、安全合规和漏洞赏金数据库。
Awaiting translation
Addy Osmani describes how he runs 5 to 10 agents in parallel every day, usually with at most 5 running at once, and breaks loop engineering into two Claude Code primitives: goal, which drives a single bounded task toward a verifiable definition of done, and loop, which reruns a prompt at fixed intervals, much like cron.
Why it matters: The author breaks his daily routine of running 5 to 10 agents in parallel into two primitives—goal and loop—and shares reusable validation Skills and ways to write stop conditions.
作者结合 Birgitta Böckeler 关于 TDD 的争论,分享了自己在 AI 智能体开发循环中的实践:当问题无法提前清晰定义时,TDD 最有用。他会像结对编程一样贴近智能体,先写一个早期测试、做小改动并观察反馈,这些测试能暴露初始需求中难以察觉的假设。等问题清晰后,他再退后让智能体完成更多实现,最后回来审查代码、运行检查并验证结果。
Awaiting translation
Augment Code has extended its Cosmos review system from code review to a full PR-to-merge loop, adding four capabilities: Verifier, PR Fixer, Review Dashboard, and cosmos approve. Dedicated Experts handle risk analysis, line-by-line correctness review, design review, runtime verification, and fixes.
Why it matters: Augment has expanded code review into a PR-to-merge loop covering fixes, verification, and approval, giving readers a way to judge how multi-agent division of labor plays out in practice.
有开发者指出,AI 只提升了写代码的速度,但正经公司里 90% 的时间花在跨部门协调、审批和排期上,效率提升非常有限。他举例称一个涉及三个部门、5 个系统的紧急需求,光沟通协调就耗掉三天,最后自己只加了三行配置。另有开发者称,20-30 人的团队如今 3 人加 5 个 Claude 200 刀账号就能同时接 20 个项目。
Awaiting translation
Martin Fowler 用 Sonnet 4.6 生成、Opus 4.8 盲评的方式,对小型、中型和较大型三类业务逻辑任务分别跑 TDD 与非 TDD 方案,结论是两者质量没有明显差异,非 TDD 方案在设计和测试质量上还多次略高,变异分数也没有实质差别。
Awaiting translation
Why it matters: 作者用同一批任务对比 TDD 与非 TDD 智能体实现,给出 token 成本与设计质量差异,并反思哪些 TDD 收益在智能体循环里失效。
作者在评审数千个编码 Agent 的改动后发现,过早加入 rescue 块、兜底逻辑和日志,往往说明功能构建顺序错了。Agent 倾向横向铺开数据库、服务、API、校验和 UI,提前设想完整系统,这与 SWE-bench 只衡量补丁能否解决限定问题并通过测试的评估方式相符。
Awaiting translation
Addy Osmani 认为,Agent 每天产生数十万甚至上百万次改动,逐行人工评审已不可行,代码质量要靠在 Agent 周围设置质量门禁来保证。这些约束包括单元测试、属性测试、验收测试、变异测试,以及圈复杂度和行长等代码质量指标,并分布在改动开始前、Agent 工作过程中和能否进入生产三个环节。
Awaiting translation
Spotify's platform team shared how the backend coding agent Honk evolved: it started out replacing migration scripts, then gradually took on build and test validation, raising merged PRs from 1000 per 3 months to 1000 per 10 days.
Why it matters: Spotify walks through Honk's evolution from script-based migration to a backend coding agent, focusing on how validation and standardization determine whether generated code can be merged.
tikalk/adlc-team-skills 发布一套 Agent Skills,通过 session_start 钩子注入约一百 token 的团队规则索引,任务匹配时再按需加载完整规则,规则存放在可 PR 评审的 team-ai-directives 仓库中。
Awaiting translation