Skip to content

#Testing/verification

0 items today
8/4Tue
8/2Sun
8/1Sat
7/29Wed
7/27Mon
7/24Fri
  1. Geoffrey Huntley · Blog34

    Geoffrey Huntley 加入 Antithesis,用形式化验证与确定性测试消除代码 slop

    Geoffrey Huntley 宣布加入 Antithesis,主张用形式化验证和确定性系统测试应对 AI 时代代码量激增带来的审查危机。他认为软件编写已被商品化,但验证与理解仍非免费,Antithesis 结合 LLM 对抗式代码审查与 pre-commit hook 驱动的语言分析器,可让人类和智能体交付可靠软件而无需掌握专门知识。

    Awaiting translation

7/22Wed
  1. Augment Code · Blog62

    What is loop engineering, and how are leading software engineering teams using it?

    Augment Code proposes loop engineering: designing agent loops that run from trigger to execution to validation to outcome, with agents handling the intermediate steps and humans stepping in only at checkpoints that require judgment. The article compares loop engineering with prompt engineering and context engineering as distinct layers, lays out five stages—trigger, execution, validation, outcome, and improvement—and describes four team-level loops already running in production: code review, ticket-to-PR, vulnerability remediation, and incident response.

    Why it matters: Augment Code breaks loop engineering into five stages—trigger, execution, validation, outcome, and improvement—and lays out four team-level loop patterns already running in production.

7/21Tue
  1. Johnny Butler · Agentic Engineering41

    AI 编程基准测不出什么:代码可维护性

    AI 编程基准只衡量模型能否在限定条件下完成改动,却测不出它留下的代码库是否更难维护——重复逻辑、耦合收紧、多余抽象都不会让当前测试失败,代价在后续扩展或排查线上问题时才显现。作者主张把可维护性放进交付流程:仓库级规范、清晰的架构边界、有意义的测试与静态分析,并审查重复、复杂度、架构漂移和模式不一致。

    Awaiting translation

7/20Mon
  1. Addy Osmani · Blog74

    Addy Osmani on software factories: the visible factory and the hidden factory, where validation is the bottleneck

    Addy Osmani proposes that a software factory has three layers—loop, harness, and factory. The factory isn’t a smarter agent; it’s multiple loops with harnesses feeding into a single review gate, with humans controlling the outer loop.

    Why it matters: The author breaks the software factory into three layers—loop, harness, and factory—and points out that validation, not generation, is the real bottleneck.

  2. Johnny Butler · Agentic Engineering38

    AI 智能体的能力是真的,但工程纪律要自己补

    当前 AI 智能体演示普遍在无限 token 预算、全新或精选代码库、无合规与运维压力的条件下运行,展示的是能力上限而非真实交付环境。AWS 副总裁 Marc Brooker 指出,智能体的机会受限于缺陷率,开发者调研中已有团队称其 AI 工具预算不可持续。作者建议只给智能体与验证能力和预算相匹配的自主权,用标准、检查项和可追踪的支出逐步放宽。

    Awaiting translation

7/17Fri
7/15Wed
  1. Johnny Butler · Agentic Engineering60

    Agent 写出的 Pull Request 比人类更好

    作者评审了一个 Agent 提交的 Pull Request,认为它又快又好,原因不在模型有多聪明,而在于质量门槛被写进了循环内部。这个 PR 在请求人工介入前就说明了改动了什么、刻意没动什么、该重点审查哪里、遵循了哪些规范,并附上了已运行的验证证据。作者由此提出,给 Agent 一个明确的完成定义和必须自证的标准,它就会为通过标准而优化,速度不是靠降低门槛换来的。

    Awaiting translation

7/13Mon
7/11Sat
  1. Habr · Kova13v80

    Using an evidence contract to constrain Claude Code's test conclusions: a QA Skill package

    A QA engineer distilled six months of experience doing web testing on Claude Code into an open-source Skill package called paranoid-qa. At its core is an evidence contract: Pass/Fail can only be based on actual artifacts like screenshots, request bodies, and logs; anything unverified gets marked Not tested; anything blocked by the environment gets marked Blocked; and forms must verify the real submitted payload.

    Why it matters: The author codified six months of QA experience into a testing Skill package for Claude Code, along with an evidence contract and failure checklist that can be reused directly.

7/10Fri
  1. Johnny Butler · Agentic Engineering49

    绿色通过从来不是你点合并的理由:AI 智能体自动合并 PR 的信任前提

    一些团队开始让 AI 智能体在 CI 全绿后自动合并自己提交的 PR,不再有人工介入。但绿色检查之所以能作为合并信号,靠的是检查运行前人类对意图、测试保护范围和系统风险点的判断,CI 只是确认而非替代这一判断。作者认为自动合并只适合低风险、边界清晰且有历史记录的改动,信任需要靠一次次带证据的改动积累,而不是打开一个开关。

    Awaiting translation

7/8Wed
  1. AI Hero · Skills Updates65

    AI Hero skills repo ships v1.1: adds /wayfinder, renames /to-spec and /to-tickets

    AI Hero's skills repo ships v1.1, renaming /to-prd to /to-spec, merging /to-plan and /to-issues into /to-tickets, and adding new Skills like /wayfinder, /research, and /prototype.

    Why it matters: The author walks through the full Skill flow from grilling to deployment and gives the migration commands for the renames, the merge, and the new /wayfinder—useful for anyone building an AI development workflow.

6/30Tue
  1. Kondasamy Jayaraman · Engineering Blog62

    形式化验证:AI 智能体尚未意识到的证明预言机

    作者认为测试只是抽样,形式化验证才能给出证明或反例,而 AI 智能体正好补上了形式化方法最贵的三块:写规格、解析反例、反复迭代。他给出 Z3、Dafny、TLA+、Alloy、Infer、Certora、Lean/Coq 的适用场景与安装方式,并指出验证器 CLI 可以像普通工具调用一样接入智能体循环,但规格写错、抽象层差异和状态空间爆炸仍是主要成本。

    Awaiting translation

6/27Sat
  1. Этихлид48

    从代理网关到本地 Opus:AI 工程面试题深度追问清单

    一份面向 AI 工程岗位的面试题追问清单,覆盖代理网关(如 OpenRouter/LiteLLM)、MCP/CLI、从零手写 Agent、多智能体编码流程、自定义 benchmark 以及本地部署约 27B-Q3_K_M.gguf 模型等方向。作者认为,这些题目的通用版本如今谁都能靠 vibe coding 一晚做出来,已失去简历价值,真正稀缺的是把任务打磨到"理想"状态的能力。

    Awaiting translation

6/26Fri
  1. Johnny Butler · Agentic Engineering41

    AI 编程智能体指令里的“Write Clean Code”只是愿望,不是可执行指令

    针对 AI 编程智能体指令文件中常见的“写整洁代码”“用 TDD”“遵循 SOLID”等表述,作者指出这些说法本身没错,但都不是可自我执行的指令。智能体主要复用代码库中已有的模式,在好坏混杂的真实代码库里,这类指令只是表达愿望。作者主张改为展示好与坏的具体样例,并让智能体证明自己实际遵循了哪些模式、在哪里应用、如何验证结果。

    Awaiting translation

6/23Tue
  1. Drew Breunig65

    Drew Breunig:问题在于提示词债,手调提示词就无法做到模型无关

    Drew Breunig 提出提示词债概念,认为用自然语言手写提示词来定义系统行为会带来三重后果:迭代变慢、团队难以读懂、应用被锁死在单一模型上。他引用 Datadog 报告称其观测到的流量中最常用的模型是 GPT-4o,并举例 Fable 的系统提示词把同一条版权规则重复了六次、Claude Code 让 Opus 七次要求在一次响应中返回多个工具调用。

    Awaiting translation

6/17Wed
6/15Mon
  1. Jesse Vincent78

    Superpowers 6 发布:构建提速最高 50%、token 花费降低最高 60%

    Superpowers 6 发布,作者称在 Anthropic 评测基准上构建耗时降低 50%、token 花费降低 60%,主要来自合并规范符合性与代码质量两个评审 agent、预先生成评审用的 diff 包让评审者少跑 git,以及调整编排器对任务所需 agent 类型的指引。

    Awaiting translation

    Why it matters: 作者用自建评测套件量化了 Superpowers 6 在构建耗时和 token 花费上的改进,并公开了实验记录与失败结论。

5/22Fri
5/19Tue
  1. Johnny Butler · Agentic Engineering34

    我绝不会为了智能体速度放弃二十年的交付实践

    面对 AI 编程智能体带来的提速压力,作者认为更快的代码生成不等于更快的软件交付,意图、验收标准、风险边界、验证、证据与评审等交付纪律不可让步。智能体只有在既有工作流内运作才能参与其中,工作越交给智能体,可评审性就越重要。这是 Software Dark Factory 背后的核心理念:智能体应强化而非取代经过验证的交付实践。

    Awaiting translation

5/18Mon
  1. The Agentic Engineer · Blog71

    AI Skill 落地难不是推广问题,而是缺少评测框架

    作者认为 AI Skill 采用率停在低位的原因不是培训不足,而是缺少能持续证明其在未测试场景下有效的评测框架。他建议从 10–15 名工程师访谈和生产日志构建 50–100 条用例集,用精确匹配和 LLM-as-judge 两种方式打分,并每季度人工标注 50 条校准裁判模型。

    Awaiting translation

5/16Sat
  1. Johnny Butler · Agentic Engineering22

    在信任自动合并 PR 之前,团队需要先补齐什么

    越来越多工程团队在讨论让 AI 智能体自动合并 PR,但真正的信任问题不在模型本身,而在于变更意图是否明确可审、验收标准是否事先定义、风险边界是否划定、验证是否真正执行、智能体权限是否受限,以及是否有可审查的决策记录。缺少这些,自动合并只是让决策变快,而非交付变快。Software Dark Factory 正尝试把意图、标准、约束、验证和权限边界内建到工作流中,让治理先于速度。

    Awaiting translation

5/13Wed
5/11Mon
5/9Sat
5/5Tue
  1. Drew Breunig62

    Agentic Coding 的 10 条经验

    Drew Breunig 整理出 agentic coding 的 10 条经验,面向刚上手 Codex、Claude Code、Pi 等智能体的人。他主张代码变便宜后应通过实现来学习、频繁重建,把投入放在端到端测试、记录意图和保持 spec 同步上,并自动化简单工作、把精力留给设计、性能、安全等难的部分。他强调代码便宜但维护、支持和安全并不便宜,智能体也会放大开发者已有的经验与品味。

    Awaiting translation

4/30Thu
4/27Mon
4/13Mon
3/24Tue
3/20Fri
  1. Terminal-Bench · News34

    Terminal-Bench 如何设计好的基准任务

    Terminal-Bench 发布任务设计指南,提出好的基准任务应具备对抗性、难度和可读性:指令像给资深工程师那样清晰直接,测试只验证结果而非实现细节,允许替代解法并防止 reward hacking。难度应来自问题本身,而非苛刻的输出格式或隐藏假设。作者建议亲自运行任务、检查容器、执行 oracle 并观察真实智能体轨迹,失败运行尤其能揭示任务是否真正困难。

    Awaiting translation

3/13Fri
  1. Ryan Lopopolo74

    智能体时代的生产函数变了:实现不再稀缺,验证才是

    作者认为,软件组织一直把人的实现时间当作稀缺投入,智能体打破了这个假设,因此策略必须可执行、验证必须随实现规模扩展。他以在大型 TypeScript 代码库开启 ESLint 的 no-await-in-loop 规则为例,发现 600 处违规,过去这需要昂贵的迁移,现在一个 PR 就能完成修复并补齐测试覆盖。

    Awaiting translation

    Why it matters: 作者以开启 ESLint 规则、迁移 600 处违规的亲身实践,说明智能体时代实现成本下降后,约束与验证为何成为新的稀缺环节。