Claude Code 将 auto mode 设为 Pro、Max 和 Team 套餐默认
Anthropic 从 8 月 14 日起把 auto mode 设为大多数 Claude Code 套餐新会话的默认设置。
Awaiting translation
Anthropic 从 8 月 14 日起把 auto mode 设为大多数 Claude Code 套餐新会话的默认设置。
Awaiting translation
Claude Code 新增跨会话消息功能,让一个会话把发现、状态或决定以纯文本发给另一个会话,Claude 通过 ListAgents 找目标、SendMessage 发送,用户无需手动调用这两个工具。
Awaiting translation
Spotify's platform team shared how the backend coding agent Honk evolved: it started out replacing migration scripts, then gradually took on build and test validation, raising merged PRs from 1000 per 3 months to 1000 per 10 days.
Why it matters: Spotify walks through Honk's evolution from script-based migration to a backend coding agent, focusing on how validation and standardization determine whether generated code can be merged.
作者结合自己用 Claude Code 生成代码的经历,指出 Vibe Coding 的安全风险几乎都源于没人读代码:智能体把第三方 API key 以明文字符串写进源码,以及只做浏览器端鉴权、后端数据接口完全无门禁。
Awaiting translation
作者认为把 ASD-STE100、Orwell 六规则、GovUK 等写作规范直接写进全局 CLAUDE.md 或 AGENTS.md,会让模型在推理时同时满足词汇约束,从而降低思考质量。
Awaiting translation
Simon Willison 把 2022 年 GPT-3 生成的游戏概念和 DALL-E 概念图丢给 Claude Code for web 里的 Claude Fable 5,让它独立做出一个可玩的 3D 浏览器游戏 Raccoon Heist,全程在手机上完成,共 7 次提交。
Awaiting translation
AI Hero Skills v1.2 发布,整套 Skill 打包为 Claude Code 插件,并为每个 SKILL.md 添加 Codex 元数据,使同一套 Skill 在 Claude Code 和 Codex 中通用。
Awaiting translation
tikalk/adlc-team-skills 发布一套 Agent Skills,通过 session_start 钩子注入约一百 token 的团队规则索引,任务匹配时再按需加载完整规则,规则存放在可 PR 评审的 team-ai-directives 仓库中。
Awaiting translation
SteveVitali 发布 agent-skills,一套与 harness 无关的 Agent Skills,旗舰 Skill implement-spec 接收 agent-ready spec 后自主完成分支、计划与测试矩阵、实现、两轮自审、与 spec 的差距分析、补齐、实时验证,最后产出 PR 和验收标准证据报告。
Awaiting translation
一篇面向 AI 辅助开发者的部署指南,建议先判断应用是纯前端还是需要常驻后端,再选择托管平台:前端可选 Vercel、Netlify、Cloudflare Pages,后端可选 Railway、Render。
Awaiting translation
作者用一周时间读完 Anthropic、OpenAI、三家超大规模云厂商、持久化层和开源项目的共十二个平台的文档,发现它们几乎都收敛到同一套架构:大脑(模型与智能体循环)、手(隔离沙箱执行生成代码)、脊柱(跨请求存活的持久状态与编排)。
Awaiting translation
Awaiting translation
Martin Alderson 表示自己第一次不再按原始智力挑选日常主力模型,而是按速度挑选,因为 Opus 4.6 级别的模型对写代码、整理研究、做幻灯片和数据库分析等日常任务已经够用。
Awaiting translation
anyaa-labs 在 GitHub 发布 agent-architect Skill,支持 Claude Code 和 Codex,用于评估和设计多智能体系统。
Awaiting translation
SimpleEnglish 是一个让大语言模型按 ASD-STE100 简化技术英语写作的 Agent Skill,兼容 Claude Code、Cursor、VS Code Copilot、OpenAI Codex、Gemini CLI 等遵循 Agent Skills 标准的工具,MIT 许可、无依赖。
Awaiting translation
Martin Fowler 用一个约 15 万行、几乎全由 Claude Code 和 Cursor 写成的 Rust 应用做实验,把 17,155 行的数据访问层按严格重构步骤拆分,每步后用全新子智能体执行同一个代表性改动并记录 token 消耗。
Awaiting translation
Why it matters: 作者用同一改动反复跑重构前后对比,量化出重构对智能体 token 消耗的实际影响,并指出节省来自文件切分而非代码总量下降。
The Terminal-Bench team releases Terminal-Bench 3.0, whose first version spans 7 domains and 74 tasks, with the strongest model passing about 34%. Building on Terminal-Bench 2.1, this release broadens task diversity and adds CI/CD, semantic versioning, and result migration to keep improving the benchmark.
Why it matters: Terminal-Bench 3.0 rebuilds the benchmark with 74 tasks and CI/CD-based versioning, so readers can see how the new benchmark separates models.
一位用 Cursor 和 Claude Code 构建并运营真实产品的开发者总结出编程提示词的四个要点:先给上下文再派任务、只交办一个窄任务、写明硬性约束、让 AI 自证结果。他以给 POST /api/submit 路由加限流为例,对比了"给 API 加限流"这类模糊提示与点名文件、复用现有 Redis 客户端、禁止新增依赖、限定 10 次/分钟并只返回 diff 不提交的写法。
Awaiting translation
Simon Willison 实测了在 Claude 和 ChatGPT 网页版接入自定义 MCP server 的步骤。
Awaiting translation
面对 GitHub 等基础设施被 AI 智能体流量压垮的现状,作者主张从 tokenmaxxing 转向 valuemaxxing,用任务完成数、节省时间和避免返工来衡量价值,而非 token 消耗量。他指出 Claude Code 会话默认 30 天后删除,导致已付费的上下文白白流失,并认为 Skill 是比 markdown 文件更好的上下文路由方式,但大规模管理 Skill 仍无解。
Awaiting translation
作者用 Cursor 和 Claude Code 实际发布过多个产品,把 AI 建应用拆成八个阶段,指出智能体擅长搭脚手架、写界面和生成 CRUD,但选技术栈、判断界面是否正确、防止脏数据、真机测试、域名 DNS 密钥和线上排错仍要人来做。
Awaiting translation
Thariq Shihipar, a member of Anthropic's engineering team, wrote up the new context engineering rules for Claude 5, saying the team has cut over 80% of the system prompt from Claude Code for models like Claude Opus 5 and Claude Fable 5, with no measurable loss on coding evals.
Why it matters: Anthropic lays out the new context engineering rules for Claude 5 and explains how to trim the system prompt, CLAUDE.md, and Skills.
Lovable 在内部搭建了一套进攻性安全程序,让 AI 智能体集群像人类攻击者一样探测系统入口,直到拿到可验证的漏洞证据。它用夺旗赛的思路做验证:把 flag 散布在基础设施和权限最高的产品界面中,不预埋任何漏洞,智能体取到 flag 就说明找到了真实入侵路径,而不是模型猜测。
Awaiting translation
Why it matters: Lovable 公开了用夺旗机制验证漏洞的内部攻防智能体编排方法,可迁移到自家安全测试流程。
Ingot 是面向个人开发者的本地优先库和 MCP server,为 Agent 的 Skill 指令提供证据门禁的变更控制。
Awaiting translation
Hugging Face disclosed a security incident that originated from a runaway agent while OpenAI was running the ExploitGym benchmark. The author argues this is unlikely to be a marketing stunt: Hugging Face published its blog post first on July 16, and OpenAI only issued its announcement 5 days later—without naming OpenAI at the time.
Why it matters: The author walks through the technical chain of the Hugging Face security incident piece by piece, and shares his take on the attack surface of autonomous agents and AI safety classifiers.
作者指出 eslint-plugin-playwright 的 recommended 配置里 no-wait-for-timeout、no-force-option、expect-expect 等关键规则只是 warning,不设 --max-warnings 0 就不会拦住合并,而 floating promise 等缺陷在没装插件时完全不可见。
Awaiting translation
作者认为托管智能体是 2026 年初各平台共同押注的方向,Anthropic、OpenAI、Google、AWS 都在推出同类产品,其核心特征是异步执行、面向生产环境的最小权限隔离和可并行扩展。
Awaiting translation
Simon Willison 用两条命令检查本机 Claude Code 是否已内置 Rust 版 Bun:strings ~/.local/bin/claude | grep -m1 'Bun v1' 输出 Bun v1.4.0(macOS arm64),另一条 grep 正则列出 563 个 .rs 源文件名。
Awaiting translation
一位独立开发者基于实际开发体验对比了 Cursor 与 Claude Code:Cursor 是 VS Code 分支的 AI 代码编辑器,适合内联编辑和 Tab 补全;Claude Code 是终端里的智能体编码工具,擅长多文件修改、测试修复循环和 git 操作。
Awaiting translation
作者 ykev 发布一份分步指南,教用户把闲置 Mac 变成 Claude Code 可完全控制、开启 computer use 的常驻机器,可从手机 Claude app 或主 Mac 经 SSH 操作。
Awaiting translation
Why it matters: 作者把闲置 Mac 改造成 Claude Code 常驻控制机,给出从 SSH、免密 sudo 到 computer use 的完整步骤,可迁移到任意两台机器。
独立开发者实测后只留下 Cursor 和 Claude Code 两款日常工具,合计每月 36 美元:Cursor Pro 16 美元负责编辑器内的 Tab 补全与行内修改,Claude Code 随 20 美元的 Claude Pro 提供,在终端处理跨多文件的智能体任务。
Awaiting translation
Claude Code 2.1.198 让 AskUserQuestion 在 60 秒无操作后自动返回“proceed anyway”,把原本阻塞的人工确认变成倒计时,2.1.200 才改为默认关闭、需在 /config 里开启。
Awaiting translation
Why it matters: 作者用二进制 diff 还原了 Claude Code 一次静默行为变更的来龙去脉,并给出关闭自动更新的可复用配置。
一位日常用 Cursor 和 Claude Code 交付线上应用的开发者解释了 AI 编程智能体的本质:它能读取代码库、修改真实文件、运行命令并自我检查,循环执行直到完成任务,而非只补全一行或给一段代码。他称这类工具在样板代码、补测试、跨文件重构和排查 bug 上表现可靠,但也会自信地做错事,需要人工复核。
Awaiting translation
Lovable 上线 agent integrations,任何公开已发布的 Lovable 应用都能添加 MCP server,从而被 ChatGPT、Claude 等 AI 工具直接调用。
Awaiting translation
Why it matters: Lovable 把已发布应用接入 MCP,读者可据此判断自家应用如何被 ChatGPT、Claude 直接调用。
作者提出 coding agent 等于模型加 harness,harness 指模型之外的所有代码、配置与执行逻辑,包括提示词、工具、上下文策略、Hook、沙箱、子智能体和反馈回路。
Awaiting translation
JetBrains 用 Claude Code 跑 SkillsBench 的 86 个真实编程任务,对比安装与不安装 Caveman 的效果,发现输出 Token 只从约 59.2 万降到 54.2 万,节省 8.5%,远低于项目宣传的 65%。
Awaiting translation
Grok 4.5 以 $6/MTok 输出价格发布,与托管版 GLM5.2 成本相近,作者认为这印证了"好够用"模型正让大量智能体任务转向低价模型。赢家是半导体与推理供应链,以及 Cursor 这类编码智能体——它们能靠廉价模型赚钱并掌握真实使用数据。输家方面作者态度矛盾:Anthropic 约 80% 收入来自 API 存在被替换风险,但前沿实验室可能改为只通过托管智能体平台提供最强模型。
Awaiting translation
A QA engineer distilled six months of experience doing web testing on Claude Code into an open-source Skill package called paranoid-qa. At its core is an evidence contract: Pass/Fail can only be based on actual artifacts like screenshots, request bodies, and logs; anything unverified gets marked Not tested; anything blocked by the environment gets marked Blocked; and forms must verify the real submitted payload.
Why it matters: The author codified six months of QA experience into a testing Skill package for Claude Code, along with an evidence contract and failure checklist that can be reused directly.
宝玉梳理了 ChatGPT 中 Chat、Work、Codex 三种模式的区别:Chat 回答问题,Work 跨应用完成知识工作并交付文档表格幻灯片,Codex 在代码仓库里完成开发任务并交付 diff、测试结果和 PR。
Awaiting translation
有用户反映 Claude Code 的 iOS 订阅突然变成了 free。讨论中提到官方兑换码有两种发放方式:下单时填写邮箱、付款成功后邮件接收兑换链接,或不填邮箱、付款后直接展示兑换链接。发帖者选择邮箱方式,因为用 stripe 收款时账号账单里会显示已投递到该邮箱,可作为真实发货凭证。
Awaiting translation