Skip to content

Expert perspectives

Experienced engineers’ and tool builders’ views and methods for software engineering in the AI era.

Latest curated items

Items 1–17 · 17 total
10/6Tue
  1. DEV Community · MCP78

    Prompt injection is a data-plane problem: move the boundary from the model to the tool call.

    The author argues that prompt injection shouldn't be solved by making models smarter; instead, just as SQL injection is handled with parameterized queries, the boundary should be drawn where the agent executes actions.

    Why it matters: The author draws an analogy between prompt injection and SQL injection, arguing for moving the boundary from the model to the tool-call layer, and lays out a practical approach with a strategy layer and separate read and write phases.

  2. DEV Community · MCP82

    Skills Are Not Tools: Why I Gave My Coding Agent 11 MCP Servers and It Fell Apart

    In February I hooked up 11 MCP Servers to my coding agent. Just the tool list in an empty session ate up 34,000 tokens, with Datadog alone contributing over two hundred tools. The agent got slower and kept picking the wrong tools.

    Why it matters: I'm writing this up because 11 MCP Servers of my own blew up my context, and it showed me that tools and knowledge belong in different containers.

10/4Sun
  1. 宝玉82

    Drawing on a podcast episode, Baoyu walks through how Lauren Tan, who works on Grok Bot at SpaceXAI, merged 2500 PRs in a single month: at night she lets the AI check and merge on its own, then spot-checks the next morning instead of reviewing each one.

    Quotedlauren@poteto

    i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful interviewer so i think the interview turned out really interesting both of our skill plugins work great together, so i recommend giving both a try and picking the best skills that suit your workflow https://www.youtube.com/watch?v=MN9dGgmLyso

    Why it matters: Using Lauren Tan's practice of merging 2500 PRs in a month, Baoyu explains that skipping individual reviews rests on validation Skills and rule constraints, and lays out the conditions under which he'd apply the same approach.

9/14Mon
  1. Addy Osmani · Blog80

    Addy Osmani on engineering methods for bringing AI agents into legacy codebases

    Addy Osmani suggests that when introducing AI agents into brownfield codebases, you should first make hidden constraints visible and make cheap changes trustworthy. He recommends dividing code into green, yellow, and red zones: green zones with solid tests let agents move in small, fast steps; yellow zones require writing characterization tests first; and red zones involving sensitive logic like authentication, billing, and permissions must have humans involved step by step. The zones are drawn by hand, and a yellow zone can only be upgraded to green once characterization tests exist and the module owner has reviewed the first batch of changes.

    Why it matters: The author turns the constraints of bringing agents into an old codebase into actionable rules—zoning, characterization tests, and migration units—and cites migration data from several companies as reference.

9/7Mon
9/5Sat
  1. Ryan Lopopolo66

    An agent platform built for inventing agents: decoupling capability interfaces from their implementations

    Author Ryan Lopopolo argues that an agent is a parameterized program built on top of a set of capabilities: models and configurations, reasoning and tool-call loops, computers, disks, context, Skills, tools, connectors, runtimes, network policies, identity, IAM, guardrails, I/O channels, and system prompts.

    Why it matters: Drawing on his experience building multiple agents, the author proposes a platform architecture that decouples capability interfaces from their implementations — a useful reference for teams building Agent platforms.

8/14Fri
  1. InfoQ · AI Coding Presentations80

    InfoQ 演讲:300 个精准 token 胜过 10 万个噪声 token,上下文工程的架构

    Baruch Sadogursky 与 Patrick Debois 在 InfoQ 演讲中用 Claude Code 现场演示:把全部项目文档塞进 CLAUDE.md 后,给接口加错误处理会因约定冲突返回 500,改用按描述懒加载的 Skill 后同一提示词通过测试。

    Awaiting translation

    Why it matters: 两位作者用现场演示拆解上下文工程的四类反模式,并给出 Skill、检索通道、外部记忆与评测的对应做法。

7/20Mon
  1. Addy Osmani · Blog74

    Addy Osmani on software factories: the visible factory and the hidden factory, where validation is the bottleneck

    Addy Osmani proposes that a software factory has three layers—loop, harness, and factory. The factory isn’t a smarter agent; it’s multiple loops with harnesses feeding into a single review gate, with humans controlling the outer loop.

    Why it matters: The author breaks the software factory into three layers—loop, harness, and factory—and points out that validation, not generation, is the real bottleneck.

7/17Fri
  1. Ryan Lopopolo71

    Code Red needs a maintenance loop: use Codex /goal to turn emergency fixes into ongoing operations

    Drawing on his experience during Stripe’s first code yellow, the author points out that after most code reds, all that’s left is a post-mortem and exhausted engineers, while the metrics go back to being unowned. He argues that a code red should leave behind a maintenance loop, and that OpenAI Codex’s /goal command can turn a one-off coding request into an ongoing objective with clear completion criteria, letting a persistent cluster of agents continuously watch metrics, generate interventions, and request human review.

    Why it matters: Based on his Stripe code yellow experience, the author proposes using Codex’s /goal to turn one-off emergency fixes into a long-term maintenance loop that can carry over to SLO governance.

5/8Fri
  1. 宝玉78

    Why the Claude Code team uses HTML instead of Markdown as the agent output format

    Claude Code team member Thariq makes the case for replacing Markdown with HTML as the output format for AI agents: HTML packs in more information, is easier to share, and supports two-way interaction, while Markdown's editing advantage stopped mattering once he switched to making changes through prompts.

    Why it matters: Claude Code team members explain why they use HTML instead of Markdown as the agent output format, and share prompts you can use as-is along with the scenarios they fit.

5/5Tue
  1. 宝玉78

    Boris Cherny: After Claude Code, writing code is turning into managing agents

    Boris Cherny, the creator of Anthropic's Claude Code, said in an interview at Sequoia AI Ascent that he went all of 2026 without writing a single line of code, merging dozens of PRs a day—150 in a single day at his peak—doing most of his work from his phone, with 5 to 10 sessions and hundreds of agents running at any given time, plus thousands more chewing through deep tasks overnight.

    Why it matters: Boris Cherny walks through Claude Code's path from incubation to a billion dollars in revenue, and lays out his calls: programming is solved, SaaS moats are being flattened, and organizational process is where the real edge is.

4/10Fri
  1. claude.dev · Anthropic Developer Blog74

    Anthropic 工程师谈 Claude Code 的工具设计:如何像智能体一样思考

    Anthropic 的 Thariq Shihipar 复盘了 Claude Code 工具设计中的取舍,核心主张是工具要贴合模型自身能力,而判断能力边界只能靠观察输出和反复实验。

    Awaiting translation

    Why it matters: Anthropic 工程师复盘 Claude Code 工具设计的取舍,给出可迁移到自建智能体的判断方法。

  2. Ryan Lopopolo65

    怎样才算把活干好:写清非功能性需求才能让 AI 智能体收敛

    Ryan Lopopolo 认为,AI 让验证问题变得明显,因为每个真实任务都依赖一个我们几乎从不写下来的问题,即怎样才算把活干好。产出和评审都涉及语气、品味、风险容忍度、打磨程度、可接受的捷径和完成标准等大量非功能性决策,过去团队靠组织设计、社交规范、招聘和入职把这些隐含规则传递给人,而模型无法走招聘流程,因此交给它的任务基本都欠规范。

    Awaiting translation

    Why it matters: 作者以在 OpenAI 做代码智能体的经历说明,非功能性需求不写下来,评审智能体就会陷入无休止的拉扯。

4/7Tue
  1. Jesse Vincent74

    如何用门禁而非规则约束 AI 智能体行为

    Jesse Vincent 在构建 Superpowers 时提出,提示词中的门禁(gate)比规则(rule)更能约束 AI 智能体:规则留有自我说服的退出路径,门禁则要求满足条件才能进入下一步。

    Awaiting translation

    Why it matters: 作者用自家智能体的实例区分规则与门禁,给出可迁移的提示词写法,帮助减少智能体跳过验证的行为。

3/13Fri
  1. Ryan Lopopolo71

    别再把代码当成最终产物:Symphony 作者谈规范驱动开发

    Ryan Lopopolo 以 Symphony 为例提出,代码不应被视为最终产物,规范才是分发物。Symphony 是一个基于 issue tracker 的智能体编排系统,先以 SPEC.md 形式分发,Elixir 参考实现只是衍生产物,README 邀请读者把 SPEC.md 交给自己的编码智能体,用任意语言重建。

    Awaiting translation

    Why it matters: 作者用 Symphony 的 spec 蒸馏循环说明代码只是可替换产物,为规范驱动开发提供了一条可迁移的路径。

  2. Ryan Lopopolo74

    智能体时代的生产函数变了:实现不再稀缺,验证才是

    作者认为,软件组织一直把人的实现时间当作稀缺投入,智能体打破了这个假设,因此策略必须可执行、验证必须随实现规模扩展。他以在大型 TypeScript 代码库开启 ESLint 的 no-await-in-loop 规则为例,发现 600 处违规,过去这需要昂贵的迁移,现在一个 PR 就能完成修复并补齐测试覆盖。

    Awaiting translation

    Why it matters: 作者以开启 ESLint 规则、迁移 600 处违规的亲身实践,说明智能体时代实现成本下降后,约束与验证为何成为新的稀缺环节。

1/30Fri