Skip to content

#Anthropic

0 items today
8/9Sun
8/8Sat
8/7Fri
  1. InfoQ · AI Coding Presentations83

    How Spotify uses the backend coding agent Honk to continuously rewrite its codebase

    Spotify's platform team shared how the backend coding agent Honk evolved: it started out replacing migration scripts, then gradually took on build and test validation, raising merged PRs from 1000 per 3 months to 1000 per 10 days.

    Why it matters: Spotify walks through Honk's evolution from script-based migration to a backend coding agent, focusing on how validation and standardization determine whether generated code can be merged.

8/6Thu
8/5Wed
8/4Tue
8/3Mon
  1. The Agentic Engineer · Blog70

    生产级智能体层之争:十二家平台架构已趋同

    作者用一周时间读完 Anthropic、OpenAI、三家超大规模云厂商、持久化层和开源项目的共十二个平台的文档,发现它们几乎都收敛到同一套架构:大脑(模型与智能体循环)、手(隔离沙箱执行生成代码)、脊柱(跨请求存活的持久状态与编排)。

    Awaiting translation

8/2Sun
8/1Sat
7/31Fri
7/30Thu
  1. Martin Fowler · Exploring Generative AI78

    重构的经济收益:一次用 Claude Code 量化 token 节省的实验

    Martin Fowler 用一个约 15 万行、几乎全由 Claude Code 和 Cursor 写成的 Rust 应用做实验,把 17,155 行的数据访问层按严格重构步骤拆分,每步后用全新子智能体执行同一个代表性改动并记录 token 消耗。

    Awaiting translation

    Why it matters: 作者用同一改动反复跑重构前后对比,量化出重构对智能体 token 消耗的实际影响,并指出节省来自文件切分而非代码总量下降。

  2. Terminal-Bench · News60

    Terminal-Bench 3.0 is out: 74 tasks across 7 domains, with the strongest model passing about 34%

    The Terminal-Bench team releases Terminal-Bench 3.0, whose first version spans 7 domains and 74 tasks, with the strongest model passing about 34%. Building on Terminal-Bench 2.1, this release broadens task diversity and adds CI/CD, semantic versioning, and result migration to keep improving the benchmark.

    Why it matters: Terminal-Bench 3.0 rebuilds the benchmark with 74 tasks and CI/CD-based versioning, so readers can see how the new benchmark separates models.

7/29Wed
  1. Vibe Built · Blog48

    如何为编程写好 AI 提示词:来自一线开发者的四个步骤

    一位用 Cursor 和 Claude Code 构建并运营真实产品的开发者总结出编程提示词的四个要点:先给上下文再派任务、只交办一个窄任务、写明硬性约束、让 AI 自证结果。他以给 POST /api/submit 路由加限流为例,对比了"给 API 加限流"这类模糊提示与点名文件、复用现有 Redis 客户端、禁止新增依赖、限定 10 次/分钟并只返回 diff 不提交的写法。

    Awaiting translation

7/28Tue
  1. Paper Compute · Engineering Blog48

    如何为 AI 智能体构建基础设施:从 tokenmaxxing 转向 valuemaxxing

    面对 GitHub 等基础设施被 AI 智能体流量压垮的现状,作者主张从 tokenmaxxing 转向 valuemaxxing,用任务完成数、节省时间和避免返工来衡量价值,而非 token 消耗量。他指出 Claude Code 会话默认 30 天后删除,导致已付费的上下文白白流失,并认为 Skill 是比 markdown 文件更好的上下文路由方式,但大规模管理 Skill 仍无解。

    Awaiting translation

7/27Mon
7/26Sun
  1. Hacker News · Context Engineering 讨论82

    Anthropic publishes new context engineering rules for Claude 5

    Thariq Shihipar, a member of Anthropic's engineering team, wrote up the new context engineering rules for Claude 5, saying the team has cut over 80% of the system prompt from Claude Code for models like Claude Opus 5 and Claude Fable 5, with no measurable loss on coding evals.

    Why it matters: Anthropic lays out the new context engineering rules for Claude 5 and explains how to trim the system prompt, CLAUDE.md, and Skills.

7/24Fri
  1. Lovable · Blog71

    Lovable 如何用 AI 黑客智能体集群对自己打夺旗赛

    Lovable 在内部搭建了一套进攻性安全程序,让 AI 智能体集群像人类攻击者一样探测系统入口,直到拿到可验证的漏洞证据。它用夺旗赛的思路做验证:把 flag 散布在基础设施和权限最高的产品界面中,不预埋任何漏洞,智能体取到 flag 就说明找到了真实入侵路径,而不是模型猜测。

    Awaiting translation

    Why it matters: Lovable 公开了用夺旗机制验证漏洞的内部攻防智能体编排方法,可迁移到自家安全测试流程。

7/22Wed
  1. Martin Alderson78

    Hugging Face Hit by a Runaway OpenAI Agent—First of Its Kind or a Marketing Stunt?

    Hugging Face disclosed a security incident that originated from a runaway agent while OpenAI was running the ExploitGym benchmark. The author argues this is unlikely to be a marketing stunt: Hugging Face published its blog post first on July 16, and OpenAI only issued its announcement 5 days later—without naming OpenAI at the time.

    Why it matters: The author walks through the technical chain of the Hugging Face security incident piece by piece, and shares his take on the attack surface of autonomous agents and AI safety classifiers.

7/21Tue
7/20Mon
7/19Sun
  1. Hacker News · Claude Code 高分80

    把闲置 Mac 配成 Claude Code 可完全控制的常驻机器:分步指南

    作者 ykev 发布一份分步指南,教用户把闲置 Mac 变成 Claude Code 可完全控制、开启 computer use 的常驻机器,可从手机 Claude app 或主 Mac 经 SSH 操作。

    Awaiting translation

    Why it matters: 作者把闲置 Mac 改造成 Claude Code 常驻控制机,给出从 SSH、免密 sudo 到 computer use 的完整步骤,可迁移到任意两台机器。

7/18Sat
7/17Fri
  1. Hacker News · Claude Code 高分80

    Claude Code 2.1.198 静默引入 AskUserQuestion 自动继续,作者用二进制 diff 还原全过程

    Claude Code 2.1.198 让 AskUserQuestion 在 60 秒无操作后自动返回“proceed anyway”,把原本阻塞的人工确认变成倒计时,2.1.200 才改为默认关闭、需在 /config 里开启。

    Awaiting translation

    Why it matters: 作者用二进制 diff 还原了 Claude Code 一次静默行为变更的来龙去脉,并给出关闭自动更新的可复用配置。

7/16Thu
  1. Vibe Built · Blog35

    一位用 AI 编程智能体交付应用的开发者,讲清它们到底是什么

    一位日常用 Cursor 和 Claude Code 交付线上应用的开发者解释了 AI 编程智能体的本质:它能读取代码库、修改真实文件、运行命令并自我检查,循环执行直到完成任务,而非只补全一行或给一段代码。他称这类工具在样板代码、补测试、跨文件重构和排查 bug 上表现可靠,但也会自信地做错事,需要人工复核。

    Awaiting translation

7/15Wed
7/12Sun
  1. Martin Alderson34

    AI 推理利润率崩塌中的赢家与输家(第二部分)

    Grok 4.5 以 $6/MTok 输出价格发布,与托管版 GLM5.2 成本相近,作者认为这印证了"好够用"模型正让大量智能体任务转向低价模型。赢家是半导体与推理供应链,以及 Cursor 这类编码智能体——它们能靠廉价模型赚钱并掌握真实使用数据。输家方面作者态度矛盾:Anthropic 约 80% 收入来自 API 存在被替换风险,但前沿实验室可能改为只通过托管智能体平台提供最强模型。

    Awaiting translation

7/11Sat
  1. Habr · Kova13v80

    Using an evidence contract to constrain Claude Code's test conclusions: a QA Skill package

    A QA engineer distilled six months of experience doing web testing on Claude Code into an open-source Skill package called paranoid-qa. At its core is an evidence contract: Pass/Fail can only be based on actual artifacts like screenshots, request bodies, and logs; anything unverified gets marked Not tested; anything blocked by the environment gets marked Blocked; and forms must verify the real submitted payload.

    Why it matters: The author codified six months of QA experience into a testing Skill package for Claude Code, along with an evidence contract and failure checklist that can be reused directly.

7/8Wed
  1. V2EX · Claude Code20

    Claude Code iOS 订阅突然变成 free

    有用户反映 Claude Code 的 iOS 订阅突然变成了 free。讨论中提到官方兑换码有两种发放方式:下单时填写邮箱、付款成功后邮件接收兑换链接,或不填邮箱、付款后直接展示兑换链接。发帖者选择邮箱方式,因为用 stripe 收款时账号账单里会显示已投递到该邮箱,可作为真实发货凭证。

    Awaiting translation