Skip to content

All updates

0 items today
12/22Mon
  1. Hacker News · AI Code Review 讨论38

    在 PR URL 后加 .diff,10 秒获得 AI 代码审查

    在任意 GitHub PR 链接末尾加上 .diff,复制原始 diff 粘贴进 Claude、ChatGPT 等 LLM,即可在 10 秒内获得初步代码审查,无需 Copilot Enterprise、浏览器扩展或特殊工具。作者强调这不能替代同事的真实代码审查,但能快速发现明显问题、补充遗漏的边界情况,缩短开发周期。

    Awaiting translation

12/20Sat
  1. Jesse Vincent70

    用 Codex 和 ChatGPT 5.2 反编译 Android 游戏 Wordiest 并移植到 iOS

    作者把 Wordiest 最后一个 Android APK 交给 Codex + ChatGPT 5.2,半小时内得到可玩的核心游戏,几小时后完成 Android 与 iOS 版本,全程未写也未读一行代码。

    Awaiting translation

    Why it matters: 作者用 Codex 反编译 Android 游戏并移植到 iOS,展示了智能体开发中难易直觉失效的真实体验。

12/19Fri
  1. Paper Compute · Engineering Blog20

    Paper Compute 为 AI 智能体构建缺失的可靠性底座

    Paper Compute 提出为 AI 智能体补上缺失的 harness,用分布式系统原语解决可靠性问题,而非依赖更强的模型。其三项能力包括确定性重放、无限上下文虚拟化和可验证状态转换,后者要求智能体的变更符合 ACID 且可审计。公司同时强调可观测性优先,认为遥测是安全、成本控制与故障恢复的前提。

    Awaiting translation

12/18Thu
12/17Wed
  1. Jesse Vincent66

    Claude Code's Skill not triggering? Maybe it never saw it at all.

    Claude Code lets the model know which Skills exist by injecting their names and descriptions into the system prompt. When there are too many Skills, or the description fields are too long, the system prompt stops listing them, so the model can't use them — and the prompt also tells the model not to use any Skill that isn't listed.

    Why it matters: The author explains why Claude Code doesn't trigger installed Skills, and gives a temporary fix using environment variables that you can apply right away.

12/15Mon
  1. Permission Protocol · AI Agent Incident Tracker60

    报道称 AI 编码机器人失误导致 AWS 服务中断

    据 Financial Times 报道,中国区一次持续 13 小时的 AWS 服务中断被归因于使用 Amazon Kiro AI 编码智能体时的用户操作失误。Amazon 据称将该事件描述为影响极其有限。报道指出,生产环境的删除、重建或发布变更本应经过签名审批流程,但公开信息未披露具体控制边界,因此无法确认智能体是否触及部署代码、内部工具或发布操作。

    Awaiting translation

12/12Fri
  1. Hacker News · AI Code Review 讨论66

    DeepSource 发布 Autofix Bot:静态分析与 AI 混合的代码审查智能体

    DeepSource(YC W20)团队发布 Autofix Bot,一个把静态分析与前沿 AI 智能体结合的代码审查智能体,面向 AI 编码工作流。其混合架构分三步:5000+ 确定性检查器建立高精度基线并由子智能体抑制误报,AI 审查以静态发现为锚点并调用 AST、数据流图、控制流、导入图等工具,最后由子智能体生成修复、静态校验后输出干净的 git patch。

    Awaiting translation

12/10Wed
  1. Jesse Vincent68

    packnplay: run a coding agent in a container with one command

    Jesse Vincent has released an open-source tool called packnplay. With a single command — `packnplay run claude --dangerously-skip-permissions` — it spins up a pre-configured throwaway container to run a coding agent.

    Why it matters: The author wraps all the tedious setup for running a coding agent in a container into one command, and also shares exactly how he handles credential conflicts with Claude Code.

12/8Mon
  1. Martin Alderson34

    构建软件的成本真的下降 90% 了吗?

    拥有近 20 年经验的开发者 Martin Alderson 认为,agentic coding 正把软件开发的劳动力成本大幅压低:原本一个月的内部工具项目现在一周完成,Claude Code 数小时就能写出 300+ 个单元与集成测试。他援引 Jevons 悖论指出,成本下降会释放大量被压抑的软件需求,而领域知识将成为开发者唯一的护城河。

    Awaiting translation

12/5Fri
  1. Simon Willison · TIL22

    pytest 9.0.0+ 新增 subtests 功能

    pytest 9.0.0 于 2025 年 11 月 8 日发布,最大新功能是内置 subtests,此前需依赖独立的 pytest-subtests 插件。subtests 作为新的默认 fixture,允许测试在运行时以编程方式动态生成子测试,不再依赖收集阶段就已知的参数列表。

    Awaiting translation

12/3Wed
12/2Tue
  1. Jesse Vincent69

    Building a front-end/back-end log bridge for the coding agent to make debugging web apps easier

    When developing web apps with the coding agent, the author often runs into client-side JavaScript bugs. If the agent can't fix them by reading the code, it fires up browser MCP for interactive debugging just to see the browser console logs—burning tokens and slowing things down.

    Why it matters: The author shares a reusable front-end/back-end log bridge approach that lets the coding agent see front-end logs without browser MCP.

12/1Mon
  1. Jesse Vincent28

    Claude Opus 4.5 谈 MCP 设计:为凌晨两点半的运维新手而设计

    Claude Opus 4.5 在阅读一篇关于 MCP 设计的文章后提出,MCP 应让用户无需理解 JMAP 的 blob 架构就能读邮件,设计目标是"凌晨两点半的运维新手"。作者用打过补丁的 JMAP MCP 服务器让 Claude 清理收件箱并学习其写作风格,随后尝试让它代为处理拖延已久的邮件,结果并不理想,如今转而让 Claude 自己构建 JMAP MCP。

    Awaiting translation

11/30Sun
11/25Tue
11/24Mon
11/21Fri
11/19Wed
  1. OpenAI · Codex Cookbook67

    How to modernize a legacy codebase in phases with Codex CLI

    In the Codex Cookbook, OpenAI lays out a complete workflow for modernizing a legacy codebase with Codex CLI, using a COBOL portfolio system as the example and moving through five phases built around an ExecPlan design document.

    Why it matters: Using a COBOL portfolio system as the example, it offers reusable documents and a validation workflow for modernizing legacy code in phases with Codex CLI.

11/16Sun
11/7Fri
  1. Terminal-Bench · News62

    Terminal-Bench ships version 2.0 and an optimized Harbor evaluation package

    Terminal-Bench has released version 2.0 and the Harbor package. The former is a more rigorously validated, harder benchmark for evaluating agents; the latter is for evaluating and optimizing agents. Harbor rewrites Terminal-Bench's test harness, supports deploying containers in the cloud, provides rollout interfaces for RL and SFT, and works with any agent you can put in a container.

    Why it matters: Terminal-Bench 2.0 and Harbor are released together, so readers can see how the agent evaluation benchmark is validated and how to scale it in the cloud.

11/2Sun
  1. Martin Alderson34

    Excel 智能体能否释放 1 万亿美元经济价值?

    继 Shortcut 和 Microsoft 的 Agent Mode 之后,Claude for Excel 的推出让 Excel 智能体成为新热点。作者估算,美国约 7090 万管理及专业岗位从业者中,38% 的工作时间花在 Excel 上,对应约 2.4 万亿美元年人力成本,即使只提升 50% 效率,也意味着至少 1 万亿美元被浪费的工时。

    Awaiting translation

10/27Mon
  1. Jesse Vincent74

    Porting Skills and Superpowers to the OpenAI Codex CLI

    Author Jesse Vincent spent an afternoon porting Superpowers and the whole SKILL.md system to the OpenAI Codex CLI, shipping it with Superpowers 3.3.0.

    Why it matters: The author ported Claude's SKILL.md system to the Codex CLI, with tool mappings and install instructions, so you can judge whether reusing Skills across models is feasible.

  2. OpenAI Developer Blog · Codex64

    Dagster Labs 如何用 Codex 做技术文档与教学

    Dagster Labs 分享了用 OpenAI Codex 加速技术文档写作、跨媒介内容转换和文档覆盖度评估的实践。他们重写了 CONTRIBUTING.md,明确文档层级、结构和最佳实践,让 Codex 能据此生成符合规范的文档;还借助 gh 命令让 Codex 解读 PR 的 diff 和描述,并让 Codex 把教程改写成 YouTube 视频脚本。

    Awaiting translation

    Why it matters: Dagster 团队把 Codex 用于文档写作、PR 解读和内容跨媒介转换,其中用文档生成代码来反向衡量文档覆盖度的做法可以迁移。

10/25Sat
  1. Jesse Vincent62

    一个让 Claude Code 派两个子智能体竞争评审代码的提示词技巧

    作者分享了一个用于 Claude Code 的代码评审提示词:让 Claude 派两个子智能体仔细评审第 5 阶段,告诉它们彼此在竞争,要求同时检查架构和实现,并称找出更多问题的一方会获得晋升。作者表示这个简单提示词的效果远超预期,并提到后续还会写更多关于代码评审提示词的内容。

    Awaiting translation

10/24Fri
  1. Jesse Vincent36

    用 ChatGPT Atlas 的 Agent 模式清理 Facebook 信息流

    作者用 ChatGPT Atlas 的 Agent 模式自动清理 Facebook 信息流,通过一段提示词让智能体持续滚动页面、隐藏赞助帖并对含 Follow/Join 链接的帖子点“不感兴趣”,最终信息流只剩自己关注的人发布的真实帖子。作者对 AI 浏览器的提示注入风险仍持警惕,表示目前不会把银行或邮箱凭据交给这类浏览器。

    Awaiting translation

10/23Thu
  1. Jesse Vincent69

    Using episodic-memory to give Claude Code cross-session memory

    The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.

    Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.

10/19Sun
  1. Jesse Vincent71

    The author built a custom superpowers-chrome MCP that cut startup overhead from 13678 tokens to 947.

    The author built a lightweight Chrome MCP and Skill for Claude Code called superpowers-chrome. At startup, the MCP configuration takes up only 947 tokens, while Microsoft's Playwright MCP needs 13678 tokens just to be available—about 7% of the context window.

    Why it matters: The author compares the token overhead of a self-built Chrome MCP against Playwright MCP, laying out the concrete trade-offs involved in designing tool interfaces for LLMs.

10/17Fri
10/16Thu
  1. Jesse Vincent78

    Anthropic launches its official Skills system across Claude Code, Claude.ai, and the Claude API

    Anthropic rolled out its first-party Skills system simultaneously on Claude Code, Claude.ai, and the Claude API, and author Jesse Vincent quickly followed with a new version of Superpowers built on the official Skills.

    Why it matters: Drawing on nearly a month of hands-on use, the author compares the official Skills with his own setup and lays out the trade-offs involved in migrating.

10/15Wed
  1. Martin Fowler · Exploring Generative AI74

    拆解 Spec-Driven Development:Kiro、spec-kit 与 Tessl 三种工具实测

    Martin Fowler 试用 Kiro、spec-kit 和 Tessl 三款自称实现 spec-driven development(SDD)的工具,把 SDD 归纳为 spec-first、spec-anchored、spec-as-source 三个层次,并指出目前所有方案都停留在 spec-first。

    Awaiting translation

    Why it matters: 作者亲手试用 Kiro、spec-kit 和 Tessl 三款 SDD 工具,给出 spec-first、spec-anchored、spec-as-source 三层划分,并指出小任务被过度规格化的问题。