Skip to content
10/6 · Tue

Latest curated items

10/4Sun
  1. DEV Community · Cursor76

    Cursor ships Composer 2, and the API response strings give away its undisclosed Kimi K2.5 base

    On March 20, 2026, developer Fynn was debugging Cursor's OpenAI-compatible endpoint when the returned model ID came back as accounts/anysphere/models/kimi-k2p5-rl-0317-s515-fast — evidence that Composer 2 was post-trained with reinforcement learning on top of Moonshot AI's Kimi K2.5. The tweet hit 44.4 views within a day.

    Why it matters: One API debugging session ties together Cursor's undisclosed Kimi base, the licensing attribution dispute, and the cost landscape for Chinese versus U.S. models — a look at how the industry handles disclosure.

9/7Mon
  1. Armin Ronacher78

    Armin Ronacher:GPT 6 Astra 写代码为什么让我不放心

    Armin Ronacher 用 GPT 6 Astra 跑了一个周末的软件工厂实验,让模型自行管理上下文和子智能体,目标是实现带虚拟线程和词法作用域的 Python。

    Awaiting translation

    Why it matters: 作者用 35 小时无人值守的软件工厂实验,展示 GPT 6 Astra 为 token 效率牺牲代码可读性的具体证据。

9/5Sat
  1. Ryan Lopopolo66

    An agent platform built for inventing agents: decoupling capability interfaces from their implementations

    Author Ryan Lopopolo argues that an agent is a parameterized program built on top of a set of capabilities: models and configurations, reasoning and tool-call loops, computers, disks, context, Skills, tools, connectors, runtimes, network policies, identity, IAM, guardrails, I/O channels, and system prompts.

    Why it matters: Drawing on his experience building multiple agents, the author proposes a platform architecture that decouples capability interfaces from their implementations — a useful reference for teams building Agent platforms.

9/2Wed
  1. Paper Compute · Engineering Blog76

    别只量代码,量工程决策:用会话记录算出一次架构决策的 65 倍放大

    作者提出用 agent 会话记录衡量一次工程决策的下游影响,即 blast radius(影响范围),并用自家 tapes 项目的一次架构决策做验证:设计文档会话花费 57.05 美元,后续引发 4704.74 美元工作量,分布在 70 个人工会话、3 名工程师、8 个仓库和 27 天中,另有 404 个自动化评测会话花费 431.54 美元。

    Awaiting translation

    Why it matters: 作者用自家一次架构决策的 474 条会话记录,展示如何把决策的下游成本量化成可复用的指标。

8/10Mon
  1. Martin Fowler · Exploring Generative AI74

    智能体循环里的 TDD 是形式还是真价值?Martin Fowler 的对比实验

    Martin Fowler 用 Sonnet 4.6 生成、Opus 4.8 盲评的方式,对小型、中型和较大型三类业务逻辑任务分别跑 TDD 与非 TDD 方案,结论是两者质量没有明显差异,非 TDD 方案在设计和测试质量上还多次略高,变异分数也没有实质差别。

    Awaiting translation

    Why it matters: 作者用同一批任务对比 TDD 与非 TDD 智能体实现,给出 token 成本与设计质量差异,并反思哪些 TDD 收益在智能体循环里失效。

7/22Wed
  1. Augment Code · Blog62

    What is loop engineering, and how are leading software engineering teams using it?

    Augment Code proposes loop engineering: designing agent loops that run from trigger to execution to validation to outcome, with agents handling the intermediate steps and humans stepping in only at checkpoints that require judgment. The article compares loop engineering with prompt engineering and context engineering as distinct layers, lays out five stages—trigger, execution, validation, outcome, and improvement—and describes four team-level loops already running in production: code review, ticket-to-PR, vulnerability remediation, and incident response.

    Why it matters: Augment Code breaks loop engineering into five stages—trigger, execution, validation, outcome, and improvement—and lays out four team-level loop patterns already running in production.

7/20Mon
  1. Addy Osmani · Blog74

    Addy Osmani on software factories: the visible factory and the hidden factory, where validation is the bottleneck

    Addy Osmani proposes that a software factory has three layers—loop, harness, and factory. The factory isn’t a smarter agent; it’s multiple loops with harnesses feeding into a single review gate, with humans controlling the outer loop.

    Why it matters: The author breaks the software factory into three layers—loop, harness, and factory—and points out that validation, not generation, is the real bottleneck.

7/17Fri
  1. Ryan Lopopolo71

    Code Red needs a maintenance loop: use Codex /goal to turn emergency fixes into ongoing operations

    Drawing on his experience during Stripe’s first code yellow, the author points out that after most code reds, all that’s left is a post-mortem and exhausted engineers, while the metrics go back to being unowned. He argues that a code red should leave behind a maintenance loop, and that OpenAI Codex’s /goal command can turn a one-off coding request into an ongoing objective with clear completion criteria, letting a persistent cluster of agents continuously watch metrics, generate interventions, and request human review.

    Why it matters: Based on his Stripe code yellow experience, the author proposes using Codex’s /goal to turn one-off emergency fixes into a long-term maintenance loop that can carry over to SLO governance.

5/12Tue
  1. Augment Code · Blog60

    Augment Code 调研 219 位工程负责人:AI 原生开发中的信任与角色落差

    Augment Code 调研了 219 位工程负责人,发现其团队约 48% 的代码由 AI 生成,55% 担心团队对代码库失去共同理解,63% 表示工程师向管理者提出了技能相关性方面的担忧,在 201-1000 人规模的团队中这一比例升至 89%。

    Awaiting translation

    Why it matters: 219 位工程负责人的调研数据揭示了 AI 原生开发中代码评审、技能焦虑与角色定义之间的落差。

5/5Tue
  1. 宝玉78

    Boris Cherny: After Claude Code, writing code is turning into managing agents

    Boris Cherny, the creator of Anthropic's Claude Code, said in an interview at Sequoia AI Ascent that he went all of 2026 without writing a single line of code, merging dozens of PRs a day—150 in a single day at his peak—doing most of his work from his phone, with 5 to 10 sessions and hundreds of agents running at any given time, plus thousands more chewing through deep tasks overnight.

    Why it matters: Boris Cherny walks through Claude Code's path from incubation to a billion dollars in revenue, and lays out his calls: programming is solved, SaaS moats are being flattened, and organizational process is where the real edge is.

4/10Fri
  1. Ryan Lopopolo65

    怎样才算把活干好:写清非功能性需求才能让 AI 智能体收敛

    Ryan Lopopolo 认为,AI 让验证问题变得明显,因为每个真实任务都依赖一个我们几乎从不写下来的问题,即怎样才算把活干好。产出和评审都涉及语气、品味、风险容忍度、打磨程度、可接受的捷径和完成标准等大量非功能性决策,过去团队靠组织设计、社交规范、招聘和入职把这些隐含规则传递给人,而模型无法走招聘流程,因此交给它的任务基本都欠规范。

    Awaiting translation

    Why it matters: 作者以在 OpenAI 做代码智能体的经历说明,非功能性需求不写下来,评审智能体就会陷入无休止的拉扯。

3/13Fri
  1. Ryan Lopopolo71

    别再把代码当成最终产物:Symphony 作者谈规范驱动开发

    Ryan Lopopolo 以 Symphony 为例提出,代码不应被视为最终产物,规范才是分发物。Symphony 是一个基于 issue tracker 的智能体编排系统,先以 SPEC.md 形式分发,Elixir 参考实现只是衍生产物,README 邀请读者把 SPEC.md 交给自己的编码智能体,用任意语言重建。

    Awaiting translation

    Why it matters: 作者用 Symphony 的 spec 蒸馏循环说明代码只是可替换产物,为规范驱动开发提供了一条可迁移的路径。

  2. Ryan Lopopolo74

    智能体时代的生产函数变了:实现不再稀缺,验证才是

    作者认为,软件组织一直把人的实现时间当作稀缺投入,智能体打破了这个假设,因此策略必须可执行、验证必须随实现规模扩展。他以在大型 TypeScript 代码库开启 ESLint 的 no-await-in-loop 规则为例,发现 600 处违规,过去这需要昂贵的迁移,现在一个 PR 就能完成修复并补齐测试覆盖。

    Awaiting translation

    Why it matters: 作者以开启 ESLint 规则、迁移 600 处违规的亲身实践,说明智能体时代实现成本下降后,约束与验证为何成为新的稀缺环节。

2/17Tue
  1. Martin Alderson76

    When AI finds a zero-day in software nobody maintains, who fixes it?

    Anthropic's red-team research shows that Claude Opus 4.6 can find 500 high-severity vulnerabilities in mature open-source projects like GhostScript and OpenSC—some of which have been sitting there for decades.

    Why it matters: Using an RCE case he reproduced himself, the author shows that once AI drives the cost of finding vulnerabilities down, unmaintained software becomes the real risk surface.

1/30Fri
  1. Jesse Vincent66

    Latent Space Engineering:用情绪化提示词让 AI 智能体进入更好状态

    Jesse Vincent 提出 Latent Space Engineering,即通过提示词把模型推入更有利于完成任务的潜在空间,而非只往上下文窗口塞事实信息。

    Awaiting translation

    Why it matters: 作者提出用情绪化提示词把模型推入更佳状态,并给出多个可迁移的实操例子。

10/15Wed
  1. Martin Fowler · Exploring Generative AI74

    拆解 Spec-Driven Development:Kiro、spec-kit 与 Tessl 三种工具实测

    Martin Fowler 试用 Kiro、spec-kit 和 Tessl 三款自称实现 spec-driven development(SDD)的工具,把 SDD 归纳为 spec-first、spec-anchored、spec-as-source 三个层次,并指出目前所有方案都停留在 spec-first。

    Awaiting translation

    Why it matters: 作者亲手试用 Kiro、spec-kit 和 Tessl 三款 SDD 工具,给出 spec-first、spec-anchored、spec-as-source 三层划分,并指出小任务被过度规格化的问题。

You've reached the end