Skip to content

#Testing/verification

0 items today
9/23Wed
  1. Lovable · Blog60

    Lovable Ships Opus 5.5: Faster Builds, Quality on Par with Opus 5

    Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.

    Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.

9/22Tue
  1. Hacker News · Coding Agent 讨论71

    Foremerge 发布:在并行编码智能体之间检测意图冲突

    Foremerge 是一个位于 git 之上的本地协调层,不改变 git 和 worktree 的工作方式。每个智能体在编辑前先发布意图和将要改动的范围及操作,第二个智能体发布时若与已有意图冲突会返回 HIGH 级别的 destructive_vs_additive 提示,检测是确定性的,不依赖评审模型读代码。

    Awaiting translation

9/21Mon
  1. Vibe Code Textbook · Articles66

    不稳定测试:重跑变绿对智能体意味着什么

    文章用概率模型算出,失败率 1% 的测试要跑 299 次才有 95% 概率被发现,而单次重跑有 99% 概率变绿,因此重跑通过几乎不提供信息。作者建议智能体循环中不要用重跑覆盖原始失败记录,并把已知不稳定测试单独标记,例如用 pytest-rerunfailures 的 @pytest.mark.flaky(reruns=n) 按测试标注,而不是全局 --reruns 3。

    Awaiting translation

9/14Mon
  1. Vibe Built · Blog71

    AI 生成代码的测试:从契约和边界出发,而不只跑通 happy path

    作者提出 AI 生成代码的测试应从任务契约出发,攻击生成代码容易忽略的边界,包括无效输入、缺失身份、错误权限、空值与最大值、重复投递、重试、并发请求、schema 兼容、部分失败和回滚,并给出一张覆盖契约、输入、身份、权限、持久化、重复、重试、并发、迁移、外部服务、可观测性和回滚的失败矩阵。

    Awaiting translation

  2. Addy Osmani · Blog80

    Addy Osmani on engineering methods for bringing AI agents into legacy codebases

    Addy Osmani suggests that when introducing AI agents into brownfield codebases, you should first make hidden constraints visible and make cheap changes trustworthy. He recommends dividing code into green, yellow, and red zones: green zones with solid tests let agents move in small, fast steps; yellow zones require writing characterization tests first; and red zones involving sensitive logic like authentication, billing, and permissions must have humans involved step by step. The zones are drawn by hand, and a yellow zone can only be upgraded to green once characterization tests exist and the module owner has reviewed the first batch of changes.

    Why it matters: The author turns the constraints of bringing agents into an old codebase into actionable rules—zoning, characterization tests, and migration units—and cites migration data from several companies as reference.

9/12Sat
  1. Simon Willison · Coding Agents0

    Boris Cherny:Claude 写的生产代码应有比人类更高的门槛

    Anthropic 的 Boris Cherny 表示,Claude 编写的生产代码应比人类编写的代码有更高门槛。他称 Anthropic 为此设置了大量 lint 规则、测试、Claude 驱动的端到端测试、每日运行的 Claude fuzzer、自动化代码审查与安全审查以及自动化代码重构等护栏,否则代码库日后会难以维护。

    Awaiting translation

9/11Fri
9/8Tue
  1. Habr · Cursor82

    How Cursor engineers merge 800+ PRs a month: using validation skills and evals to build trust in AI agents

    Cursor engineer Lauren Tan shares how she gets AI agents to submit and merge PRs on their own: the key is validation—letting the agent run code, capture CPU traces, and open an iOS simulator to check its own work.

    Why it matters: Cursor engineers break trust in AI agents down into reusable validation skills, feature maps, and evals, so readers can build their own automated validation workflows.

9/6Sun
  1. Johnny Butler · Agentic Engineering60

    工程规范应放进智能体编码循环,而不是等 CI 才发现问题

    作者 Johnny Butler 认为,如果编码智能体直到 CI 阶段才发现自己违反了工程规范,反馈就来得太晚,应在智能体实现过程中就运行检查。他建议把测试、linter、类型检查、安全扫描以及团队自定义的依赖规则和复杂度阈值检查交给智能体,让它改完小改动后运行相关检查、修复失败项再重跑,并把结果随代码一起返回,使开发者能用同样的检查复现。

    Awaiting translation

9/5Sat
  1. Vibe Code Textbook · Articles78

    如何写出编码智能体能完成的任务:六段式 spec 模板与 linter

    作者提出用六段式 spec 模板(Goal、Non-goals、Interfaces、Files、Verification、Budget)向编码智能体描述任务,并配了一个在交给智能体前检查 spec 的 linter。

    Awaiting translation

    Why it matters: 给出可直接套用的六段式 spec 模板、tally 实例和配套 linter,读者能据此改造自己交给编码智能体的任务描述。

  2. Vibe Code Textbook · Articles62

    如何用一段式 spec 让 Claude Code 构建 CSV 命令行工具

    作者以 tally 这个 CSV 列统计 CLI 为例,给出从一段式 spec 到可运行工具的完整方法:spec 包含目标、非目标、接口、文件、验证和预算六部分,提示词按先列边界情况、再写九条测试、然后实现、跑三条手工检查、最后按 diff 清单评审的顺序发送。

    Awaiting translation

  3. Vibe Code Textbook · Articles66

    SWE-bench 分数到底测了什么:resolve rate 的生成过程与五篇论文的补充发现

    SWE-bench 的百分比是智能体编程领域被引用最多的数字,但它比通常被引用的方式要窄得多。文章拆解了一次评测如何产出这个数字:提交是 JSONL 格式的补丁,在 Docker 容器中应用到 base_commit 并运行仓库测试,FAIL_TO_PASS 全部通过且 PASS_TO_PASS 不回归才算 resolved,默认是 pass@1。

    Awaiting translation

  4. Vibe Code Textbook · Articles87

    审查 coding agent 的 diff:检查清单最先抓到什么

    作者给出一份按危害排序的十项 agent diff 检查清单,依次看被删除的测试、被跳过或弱化的断言、宽泛异常捕获、新增依赖、任务范围外文件、CI 配置改动、疑似密钥、遗留标记和净删除超过 40 行的文件。

    Awaiting translation

    Why it matters: 给出按危害排序的十项 agent diff 检查清单,并附可复用的扫描脚本与行号定位。

  5. Vibe Code Textbook · Articles71

    让智能体先写测试:为什么测试必须先失败再通过

    文章主张让 AI 智能体在写实现前先展示一次失败的测试,因为一次都没红过的测试无法证明它能检测任何行为。作者引用 AgentLens 论文(arXiv:2605.12925)对 2,614 条 OpenHands 轨迹的评测,发现通过轨迹中有 10.7% 属于 Lucky Pass,各模型的幸运通过率在 0.5% 到 23.2% 之间。

    Awaiting translation

9/4Fri
  1. OpenAI Developer Blog · Codex71

    How to Build a Game with Astra in Codex: From Void Explorer to Performance Tuning

    The author built the space exploration game Void Explorer in Codex with Astra, featuring 2,048 star systems and over 10,000 procedurally generated planets, and shared the full workflow from prompts to architecture, testing, and performance measurement.

    Why it matters: Using Astra in Codex, the author built an entire game and showed a transferable collaborative workflow that spans prompts, testing, and performance measurement.

9/2Wed
  1. Hacker News · Agent Skills78

    mattpocock releases AI coding Agent Skills built for real engineering

    Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.

    Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.

9/1Tue
8/27Thu
8/21Fri
  1. Sean Goedecke · Blog34

    读者无法分辨带水印的 AI 文本:一项基于 SynthID-Text 的测试

    博主用 Qwen3-30B-A3B-Instruct-2507 在租用的 H200 上生成 30 条回答,其中部分用 SynthID-Text 加水印,让读者分辨哪条被水印。首轮 278 人平均得分 3.92/10,重排题目后 73 人平均 3.4/10,接近纯随机猜测的 3.33,说明读者无法识别水印。目前测验累计约 4700 份回答,均值 3.44/10。

    Awaiting translation

  2. OpenAI Developer Blog · Codex65

    OpenAI Releases Daybreak and Codex Security, a Security Workflow

    OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.

    Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.

8/20Thu
8/18Tue
  1. Lovable · Blog74

    Lovable 如何把 lovable.dev 从 Next.js 迁到自家 TanStack Start 栈

    Lovable 用六个月把月访问 4200 万的 lovable.dev 从 Next.js 迁到自家 TanStack Start 托管栈,迁移期间两套框架并行运行,由代理 worker 按路由和用户分流,最终 Next.js 专属代码只占 3%。

    Awaiting translation

    Why it matters: Lovable 官方复盘把 90 万行代码从 Next.js 迁到自家 TanStack Start 栈,双框架并行、AI 智能体批量迁移和一次 OOM 事故的细节都可直接借鉴。

8/17Mon
8/14Fri
  1. Addy Osmani · Blog78

    Addy Osmani’s loop engineering in practice: managing parallel agents with Claude Code’s goal and loop primitives

    Addy Osmani describes how he runs 5 to 10 agents in parallel every day, usually with at most 5 running at once, and breaks loop engineering into two Claude Code primitives: goal, which drives a single bounded task toward a verifiable definition of done, and loop, which reruns a prompt at fixed intervals, much like cron.

    Why it matters: The author breaks his daily routine of running 5 to 10 agents in parallel into two primitives—goal and loop—and shares reusable validation Skills and ways to write stop conditions.

  2. Johnny Butler · Agentic Engineering36

    在 AI 智能体开发循环中,TDD 在哪些环节真正帮到我

    作者结合 Birgitta Böckeler 关于 TDD 的争论,分享了自己在 AI 智能体开发循环中的实践:当问题无法提前清晰定义时,TDD 最有用。他会像结对编程一样贴近智能体,先写一个早期测试、做小改动并观察反馈,这些测试能暴露初始需求中难以察觉的假设。等问题清晰后,他再退后让智能体完成更多实现,最后回来审查代码、运行检查并验证结果。

    Awaiting translation

8/13Thu
  1. Augment Code · Blog62

    Augment Code Expands Cosmos: Turning Code Review into an Agentic PR-to-Merge Loop

    Augment Code has extended its Cosmos review system from code review to a full PR-to-merge loop, adding four capabilities: Verifier, PR Fixer, Review Dashboard, and cosmos approve. Dedicated Experts handle risk analysis, line-by-line correctness review, design review, runtime verification, and fixes.

    Why it matters: Augment has expanded code review into a PR-to-merge loop covering fixes, verification, and approval, giving readers a way to judge how multi-agent division of labor plays out in practice.

8/10Mon
  1. V2EX · Vibe Coding34

    AI 对研发效率提升到底有多大?开发者称写代码只占工作 10%

    有开发者指出,AI 只提升了写代码的速度,但正经公司里 90% 的时间花在跨部门协调、审批和排期上,效率提升非常有限。他举例称一个涉及三个部门、5 个系统的紧急需求,光沟通协调就耗掉三天,最后自己只加了三行配置。另有开发者称,20-30 人的团队如今 3 人加 5 个 Claude 200 刀账号就能同时接 20 个项目。

    Awaiting translation

  2. Martin Fowler · Exploring Generative AI74

    智能体循环里的 TDD 是形式还是真价值?Martin Fowler 的对比实验

    Martin Fowler 用 Sonnet 4.6 生成、Opus 4.8 盲评的方式,对小型、中型和较大型三类业务逻辑任务分别跑 TDD 与非 TDD 方案,结论是两者质量没有明显差异,非 TDD 方案在设计和测试质量上还多次略高,变异分数也没有实质差别。

    Awaiting translation

    Why it matters: 作者用同一批任务对比 TDD 与非 TDD 智能体实现,给出 token 成本与设计质量差异,并反思哪些 TDD 收益在智能体循环里失效。

8/8Sat
  1. Johnny Butler · Agentic Engineering64

    过早的错误处理暴露了 Agent 把功能建错了顺序

    作者在评审数千个编码 Agent 的改动后发现,过早加入 rescue 块、兜底逻辑和日志,往往说明功能构建顺序错了。Agent 倾向横向铺开数据库、服务、API、校验和 UI,提前设想完整系统,这与 SWE-bench 只衡量补丁能否解决限定问题并通过测试的评估方式相符。

    Awaiting translation

  2. Addy Osmani · Blog62

    Addy Osmani:智能体时代的代码质量取决于你给 Agent 设的约束

    Addy Osmani 认为,Agent 每天产生数十万甚至上百万次改动,逐行人工评审已不可行,代码质量要靠在 Agent 周围设置质量门禁来保证。这些约束包括单元测试、属性测试、验收测试、变异测试,以及圈复杂度和行长等代码质量指标,并分布在改动开始前、Agent 工作过程中和能否进入生产三个环节。

    Awaiting translation

8/7Fri
  1. InfoQ · AI Coding Presentations83

    How Spotify uses the backend coding agent Honk to continuously rewrite its codebase

    Spotify's platform team shared how the backend coding agent Honk evolved: it started out replacing migration scripts, then gradually took on build and test validation, raising merged PRs from 1000 per 3 months to 1000 per 10 days.

    Why it matters: Spotify walks through Honk's evolution from script-based migration to a backend coding agent, focusing on how validation and standardization determine whether generated code can be merged.

8/4Tue