Skip to content

All updates

0 items today
10/13Mon
  1. Jesse Vincent37

    让 AI 写出好文风的一个怪招:先让 Claude 读《The Elements of Style》

    开发者 Jesse Vincent 发现,让 Claude 先读 Strunk 1920 年版《The Elements of Style》再写 README,成稿比原来短约 30%,文风也更合他意。他把该书 HTML 转成约 12,000 词的 Markdown,因 Anthropic 的版权过滤机制拒绝处理这本已进入公有领域的书,最终改用 GPT-5 Codex 完成删减。

    Awaiting translation

10/12Sun
  1. Martin Alderson40

    Martin Alderson 推出 MCP Server 增长追踪器

    Martin Alderson 搭建了一个追踪器,监控官方 MCP server registry 中 MCP 服务器的增长情况。数据显示,过去一周 MCP 服务器从约 900 个增至超过 1,150 个,每天新增约 25-30 个。该追踪器每日更新,可在 mcp-tracker.martinalderson.com 实时查看生态变化。

    Awaiting translation

10/10Fri
10/9Thu
  1. Jesse Vincent78

    Superpowers: How the Author Used a Coding Agent in October 2025

    Author Jesse Vincent released Superpowers, a set of Skills built on Claude Code's new plugin system. Once installed, it injects a guiding prompt through the session-start hook, prompting Claude to proactively search for and use these Skills.

    Why it matters: The author packaged his own coding-agent workflow into an installable Skill plugin, so readers can directly reuse his implementation flow from brainstorming to TDD.

10/5Sun
  1. Jesse Vincent71

    How I Used a Coding Agent in September 2025: A Dual-Session Workflow with Claude Code

    The author walks through their full workflow with Claude Code: first isolating tasks with git worktree, then using a brainstorming prompt to make Claude ask only one question at a time and confirm the design in stages, and finally using a planning prompt to break the plan into small tasks and write them into docs/plans/.

    Why it matters: The author splits Claude Code into two sessions—an architect and an implementer—and shares reusable prompts plus a git worktree approach for isolating tasks.

10/2Thu
10/1Wed
  1. Hacker News · Context Engineering 讨论76

    Anthropic on Effective Context Engineering for AI Agents

    Anthropic's applied AI team argues that context engineering is a continuation of prompt engineering, and the core idea is picking the smallest set of high-signal tokens within a limited attention budget.

    Why it matters: Anthropic lays out a systematic approach to context engineering, covering the trade-offs among three long-task strategies: compression, note-taking, and sub-agents.

9/30Tue
9/29Mon
  1. Jesse Vincent66

    Rewriting CLAUDE.md Rules with GraphViz's Dot Language

    The author rewrote a long block of CLAUDE.md rules as a GraphViz dot flowchart, using quoted strings as node names, different shapes to distinguish decisions, commands, and warnings, and giving each flow an explicit trigger condition.

    Why it matters: After rewriting the CLAUDE.md rules as a GraphViz dot flowchart, Claude followed the rules better, and this approach can be carried over to your own projects.

9/28Sun
9/25Thu
9/24Wed
  1. Hacker News · Context Engineering 讨论87

    Manus's AI Agent Context Engineering Experience: Six Practical Principles

    The Manus team shares context engineering lessons from building AI agents, centered on designing around the KV cache, managing tools by masking rather than removing them, treating the file system as context, steering attention by restating to-do items, keeping errors in context, and avoiding getting stuck on few-shot examples.

    Why it matters: The Manus team distilled lessons from rewriting their agent framework four times into six context engineering principles—ready to apply directly to your own agent implementation.

9/9Tue
  1. Terminal-Bench · News38

    Terminal-Bench 排行榜回应超时争议:OB-1 重新提交成绩后重回榜首

    OpenBlock Labs 的智能体 OB-1 在按正确超时限制重新提交成绩后,重新成为 Terminal-Bench 排行榜得分最高的智能体。此前该提交将数据集中每个任务的超时统一改为固定 30 分钟,而多数任务原本限时 5 分钟,导致 75 个任务超时被放宽、2 个不变、3 个被缩短。

    Awaiting translation

8/29Fri
  1. OpenAI · Codex Cookbook74

    在 GitLab CI/CD 中用 Codex CLI 自动做代码质量检查与安全修复

    OpenAI Codex Cookbook 给出把 Codex CLI 接入 GitLab CI/CD 的完整做法,用于生成 CodeClimate JSON 代码质量报告、把 SAST 结果整理成 security_priority.md,并让 Codex 输出可 git apply 的补丁。

    Awaiting translation

    Why it matters: 官方 Cookbook 给出把 Codex CLI 接入 GitLab CI 的完整配置,含提示词约束、JSON 标记提取与 diff 校验,可直接照搬。

8/26Tue
8/10Sun
  1. Ryan Lopopolo38

    MCP 如何为 LLM 解决工具发现问题

    MCP 通过提供一份实时、机器可读的工具目录,为 Claude Code、OpenAI Codex 等编程智能体解决工具发现问题——目录包含工具名称、描述、输入 schema 和示例调用。CLI 工具对编程智能体不可读,而 MCP 内置了自动提示模型的机制,让模型无需猜测或搜寻工具。作者认为,智能体优先的开发意味着从工具编写之初就要面向模型提示,这正是 MCP 所内建的。

    Awaiting translation

8/3Sun
  1. Ryan Lopopolo22

    服务网格是组织工具,而非技术工具

    服务网格的价值在于以难以绕开的方式注入策略,包括连接池与 keep alive、重试、超时、数据本地性约束和 AZ 本地路由等。产品工程追求快速迭代与可靠性之间存在一定冲突,而服务网格难以被绕过,能确保默认情况下做正确的事。这也是部署服务网格最简洁的理由:它实现了职责分离,又不需要产品团队额外记住步骤。

    Awaiting translation

8/1Fri
7/18Fri
7/15Tue
7/13Sun
7/9Wed
  1. Hacker News · Context Engineering 讨论62

    上下文工程实战:用 n8n 搭建多智能体深度研究工作流

    作者以自己用 n8n 搭建的多智能体深度研究应用为例,说明上下文工程不只是写提示词,而是设计并优化提供给 LLM 的完整上下文。他拆解了搜索规划智能体的系统提示词,涵盖指令、用户输入分隔符、子任务字段定义、结构化输出示例,以及注入当前日期时间让模型推断 start_date 和 end_date 的做法。

    Awaiting translation

7/1Tue
  1. Hacker News · Context Engineering 讨论78

    Context Engineering for Agents: Four Strategies—Write, Select, Compress, and Isolate

    In his article, Lance Martin groups context engineering for agents into four strategies: writing (using scratchpads and memory to store information outside the context window) and selecting (pulling in memory, tool descriptions, and knowledge on demand).

    Why it matters: The article groups agent context management into four strategies—writing, selecting, compressing, and isolating—and shows how various products put them into practice, making it easy to compare against your existing workflow.

  2. Hacker News · Context Engineering 讨论62

    AI 的新技能不是提示词工程,而是上下文工程

    作者认为 AI 领域正从提示词工程转向上下文工程,即设计动态系统,在合适的时间以合适的格式为 LLM 提供正确的信息和工具。他把上下文拆成系统提示词、用户提示词、短期状态与历史、长期记忆、RAG 检索信息、可用工具和结构化输出七类,并指出多数智能体失败已不是模型失败而是上下文失败。

    Awaiting translation

6/26Thu
6/25Wed
6/24Tue
6/20Fri
6/4Wed
  1. Martin Fowler · Exploring Generative AI66

    自主编码智能体实测:一次 OpenAI Codex 运行记录

    Martin Fowler 给 OpenAI Codex 布置了一个前端标签格式化的化妆类小任务,并完整公开了 Codex 的日志和生成的 PR。日志显示 Codex 主要靠 grep 反复文本搜索定位代码,中途因把 AGENTS.md 误写成 AGENT.md 来回折腾,还因删掉 .yarnrc 导致测试无法运行,最终 PR 里有两个回归测试失败。

    Awaiting translation

    Why it matters: 作者完整记录 Codex 自主完成一次前端小任务的日志,并对比 6 次运行结果,展示后台编码智能体在环境配置和代码复用上的真实短板。

5/28Wed
5/23Fri
  1. Terminal-Bench · News32

    Anthropic 在 Claude 4 模型卡中引入 Terminal-Bench,Claude 4 Opus 创下 43.2% 新 SOTA

    Anthropic 将 Terminal-Bench 列为 Claude 4 模型卡七项基准之一,Claude 4 Opus 在 Terminal-Bench-Core 上取得 43.2% 的 SOTA 成绩。Dario Amodei 在 Code with Claude 主题演讲中也提及该基准。Terminal-Bench 团队表示将在未来几天验证 Claude 4 的表现并更新官方排行榜。

    Awaiting translation

5/22Thu
  1. Jesse Vincent22

    为 Vibe Coding 做一个“continue”按钮键盘

    Jesse Vincent 用 Raspberry Pi Pico 和一只廉价 Staples Easy Button 改装出专为 AI 编程助手设计的单键键盘,按下即向 Claude Code、Cursor、Cline 等工具发送 "continue",让 AI 接着干活。他还在按钮里保留了原装扬声器播放励志语音,但指出音效很快就让人厌烦,取出电池后键盘功能仍可正常使用。

    Awaiting translation

5/19Mon
  1. Terminal-Bench · News48

    Terminal-Bench 推出研究预览版智能体 Terminus

    Terminal-Bench 团队发布研究预览版智能体 Terminus,用于在终端中一致地评估语言模型驱动自主智能体的能力,发布时其性能在 Terminal-Bench 上仅次于 Claude Code。Terminus 采用单工具设计,仅通过 tmux 会话发送标准按键操作,并借助 LiteLLM 支持几乎所有 API 或本地托管模型,且完全自主运行、不请求用户输入。

    Awaiting translation

  2. Terminal-Bench · News62

    Terminal-Bench 发布首个终端智能体评测基准

    Terminal-Bench 发布首个版本,用于量化 AI 智能体在终端中执行复杂任务的能力,首发数据集 Terminal-Bench-Core-v0 包含 80 个手工编写并人工验证的任务,每个任务配有独立 Docker 环境、人工验证的解法与测试用例。

    Awaiting translation

    Why it matters: Terminal-Bench 给出 80 个带 Docker 环境和测试用例的终端任务,可用来横向比较不同智能体在命令行中的实际表现。

5/14Wed
  1. Martin Fowler · Exploring Generative AI62

    用 LLM 构建 PlantUML 分步播放扩展 PlantUMLSteps

    作者用 LLM 构建了 PlantUMLSteps,为 PlantUML 时序图增加按步骤播放功能,把单张复杂图拆成登录、认证、仪表盘等逐步展示的步骤。开发中 Claude 在 Cursor Agent 模式下生成 StepParser 解析逻辑、Gradle 任务和 HTML 查看器,解析器最初在 newPage 属性、首个步骤标记前的声明处理上出错,经测试反馈修正后通过。

    Awaiting translation