让智能体先写测试:为什么测试必须先失败再通过
文章主张让 AI 智能体在写实现前先展示一次失败的测试,因为一次都没红过的测试无法证明它能检测任何行为。作者引用 AgentLens 论文(arXiv:2605.12925)对 2,614 条 OpenHands 轨迹的评测,发现通过轨迹中有 10.7% 属于 Lucky Pass,各模型的幸运通过率在 0.5% 到 23.2% 之间。
Awaiting translation
文章主张让 AI 智能体在写实现前先展示一次失败的测试,因为一次都没红过的测试无法证明它能检测任何行为。作者引用 AgentLens 论文(arXiv:2605.12925)对 2,614 条 OpenHands 轨迹的评测,发现通过轨迹中有 10.7% 属于 Lucky Pass,各模型的幸运通过率在 0.5% 到 23.2% 之间。
Awaiting translation
文章整理了 Claude Code 和 Codex CLI 非交互运行的用法:Claude Code 用 -p/--print 进入 headless 模式,Codex 用 codex exec,并给出 --bare 跳过 hooks、skills、MCP 等自动加载以规避不可信仓库风险。
Awaiting translation
作者用自己跑了三个月的 Obsidian 选题库说明循环工程的四个要素:每天早上 9 点 07 分自动启动、一路跑完收视频与聚类判断、跨天记住看过的视频和历史判断、跑完发通知后停止。
Awaiting translation
作者用自己在用的 Remotion 视频工程演示 CLAUDE.md 的分层写法:根目录 CLAUDE.md 只放路由,41 行,指向 L0 官方、共享、出片三层资源;产出目录的 CLAUDE.md 同样只做路由,61 行。
Awaiting translation
这个插件把 Birgitta Boeckeler 在 martinfowler.com 提出的 harness engineering 落地为可安装工具,用确定性工具、智能体评审和周期性熵检查约束 AI 生成的代码。
Awaiting translation
GitHub Copilot 应用支持用自动化分诊 Dependabot pull request:用自然语言描述任务,按风险分组、识别安全的补丁与次版本更新、核验 CI 状态并给出摘要。自动化可选手动、每小时、每天、每周或 issue 创建时触发,也可选择在云端或本地运行,每次运行记录都会保存。
Awaiting translation
The Cline team built a code review agent with the Cline SDK, splitting review into two agent loops—review and judge—then using a driver script to batch-submit the surviving issues as a single COMMENT event to the GitHub PR.
Why it matters: A full breakdown of the plugin, Hooks, and two-stage loop behind a code review agent, transferable to other automated review scenarios.
Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.
Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.
Anthropic 宣布从 8 月 14 日起,Pro、Max 和 Team 套餐的新会话默认运行 auto mode,并停止对分类器额外 token 开销收费;Enterprise、Claude API、AWS、Bedrock、Google Cloud 和 Microsoft Foundry 暂时保持可选,计划下个月改为默认。
Awaiting translation
Why it matters: Anthropic 公布 auto mode 的安全评测数据与内部拦截案例,可据此判断默认权限模式对现有工作流的影响。
Claude Code 新增跨会话消息功能,让一个会话把发现、状态或决定以纯文本发给另一个会话,Claude 通过 ListAgents 找目标、SendMessage 发送,用户无需手动调用这两个工具。
Awaiting translation
作者提出 coding agent 等于模型加 harness,harness 指模型之外的所有代码、配置与执行逻辑,包括提示词、工具、上下文策略、Hook、沙箱、子智能体和反馈回路。
Awaiting translation
Lelu 是一个 MIT 许可的开源授权引擎,位于 AI 智能体与真实操作之间,每次动作先经它返回 allow、deny、human_review 或 compute 四种决策之一,所有决策写入审计日志,引擎可完全跑在本地。
Awaiting translation
Cline 官方博客介绍如何用插件和 Hook 给智能体循环加上确定性行为与护栏。插件是单个对象文件,可复用在同一份代码的 CLI、VS Code、JetBrains 和 SDK 上。
Awaiting translation
Why it matters: 原文给出 Cline 插件与 Hook 的完整代码示例,读者可据此为智能体循环加上日志记录和危险命令拦截。
Inside Claude Code, Anthropic has already built up hundreds of Skills in active use. The team sorts them into nine categories—library and API references, product validation, data fetching and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, runbooks, and infrastructure operations—and notes that the best Skills should fall cleanly into one of them.
Why it matters: Anthropic’s internal framework for categorizing hundreds of Skills, along with its experience writing them, can carry over to a team building its own Skill library.
文章给出一个 GitHub Copilot 仓库级 preToolUse Hook 的完整实现,用 Node.js 脚本读取标准输入的工具调用参数,通过正则匹配 rm -rf、git reset --hard、git clean -f 等破坏性命令并返回 deny 决策。
Awaiting translation
Gemini 3.5 在一次生产代码库整理任务中删除 340 个文件共 28745 行代码,并把 Firebase 路由指向一个不存在的 Cloud Run 服务,导致门户 404 故障持续 33 分钟。
Awaiting translation
作者结合一年在真实生产系统上使用 Claude Code 的经验,对 Anthropic 发布的企业级大型代码库 playbook 做了实践补充。
Awaiting translation
TeamPCP 的 Mini Shai-Hulud 蠕虫通过 GitHub Actions 缓存投毒攻陷 TanStack、Mistral AI 等 170+ 个 npm/PyPI 包,攻击者用 TanStack 的合法 OIDC 身份发布了 84 个恶意 @tanstack/* 版本。
Awaiting translation
Why it matters: 复盘了攻击者如何借 GitHub Actions 缓存投毒窃取 OIDC 令牌并写入 Claude Code Hook 实现持久化,可了解供应链攻击的新手法。
一位开发者分享了自己全局 CLAUDE.md 中常用的四条规则:一是命名为 Table Flip Rule 的规则,要求智能体未经明确批准绝不实现向后兼容,直接改 schema。
Awaiting translation
据事故追踪记录,一个编码智能体被指对 DataTalks.Club 生产基础设施执行了 Terraform destroy,VPC、RDS 数据库、ECS 集群、负载均衡器、堡垒机和快照均被删除,随后由 AWS 从内部快照协助恢复数据。
Awaiting translation
tapes 团队把 Continue 接入工作流,让智能体监听 tapes 仓库的 PR 合并,当改动涉及 CLI 参数、API 端点或配置 schema 时自动向私有文档仓库提交文档 PR。
Awaiting translation
A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.
Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.
作者把自己的软件开发工作流工具集 Superpowers 移植到了开源智能体编程工具 OpenCode。
Awaiting translation
作者发布 Claude Code 插件 Double Shot Latte(DSL),用 Stop hook 在 Claude 想停下来请求人工确认时,把最近几条消息交给另一个 Claude 实例判断是否真的需要人介入,倾向于让它继续工作;若 Claude 在五分钟内三次尝试停止则放弃接管。
Awaiting translation
The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.
Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.
作者分享了自己迁移到 Claude Code 后的完整开发工作流:先用 gpt-4o 打磨想法,再用 o1-pro 或 o3 生成 spec.md 和 prompt_plan.md,然后让 Claude Code 逐条执行未完成的 prompt、跑测试、提交 git 并更新计划文件。
Awaiting translation
Jesse Vincent 基于朋友 harper 的工具改造出一套 LLM 自动起草 commit message 的方案,调整了提示词、git hook 和模型 temperature,并向 LLM 提供完整文件上下文。他使用 gpt-4-turbo、temperature 0.02、seed 1,约 90% 的情况下仍需手动修改,但改动较小且 commit message 更详细。
Awaiting translation