Skip to content

All updates

0 items today
8/25Tue
  1. Lovable · Blog62

    How Lovable connected its own app to external tech stacks: from MCP to nearly 100 connectors

    Lovable shared a retrospective on how it connected its platform app to third-party services: first it supported MCP as a stopgap for pulling context into chats, then it built app connectors of its own, using a Connector Gateway to proxy requests between published apps and third-party APIs. The gateway holds credentials and refresh logic, so deployed apps never touch the keys.

    Why it matters: Lovable’s retrospective on turning connectors into reusable infrastructure is worth a look for teams doing third-party integrations and credential management.

  2. OpenAI Developer Blog · Codex62

    Automating OpenAI’s repetitive evaluation work with Codex and the Runme notebook

    OpenAI engineers use Codex with the open-source notebook app Runme to automate repetitive work such as running model evaluations. The approach: write a goal cell in the Runme notebook, have Codex read the goal, produce a plan, and wait for human approval before executing, logging commands, outputs, and conclusions along the way—including the dead ends.

    Why it matters: The author uses the Runme notebook plus WebMCP to hand the evaluation process over to Codex; readers can borrow the way it handles goals, approvals, and context capture.

8/23Sun
  1. Thorsten Ball · Register Spill12

    Joy & Curiosity #96:Stripe 以 75 亿美元收购 OpenRouter,GitHub 月提交量增至 29 亿

    Stripe 以 75 亿美元收购 OpenRouter,致投资人的信被曝光。GitHub 披露月提交量自 4 月以来从 14 亿增至 29 亿,新增 300 万 CPU 核心和 120 PB 高速存储。作者还回顾了 GitHub 开源贡献、StackOverflow、TDD、文本编辑器等曾被视为基石的开发实践在近两年的式微。

    Awaiting translation

8/21Fri
  1. OpenAI Developer Blog · Codex65

    OpenAI Releases Daybreak and Codex Security, a Security Workflow

    OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.

    Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.

8/20Thu
8/19Wed
  1. Cursor · Changelog66

    Cursor Updates Cloud Agents and Harness: Event Subscriptions, Independent Sub-Agent Runs, and /goal

    Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.

    Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.

  2. Cline · Blog71

    Cline's Evaluation Methodology and Trace for Open-Weight Models

    Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.

    Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.

  3. OpenAI Developer Blog · Codex71

    OpenAI open-sources the Codex harness and the app-server client protocol

    OpenAI has open-sourced the harness that drives the Codex app, CLI, and IDE extensions, and through the Codex app-server client protocol it exposes capabilities like creating threads, starting turns, receiving events, and handling approval requests.

    Why it matters: With the Codex harness and app-server protocol now public, developers can see how to embed the agent in their own products and where the boundaries are.

8/18Tue
  1. Lovable · Blog74

    Lovable 如何把 lovable.dev 从 Next.js 迁到自家 TanStack Start 栈

    Lovable 用六个月把月访问 4200 万的 lovable.dev 从 Next.js 迁到自家 TanStack Start 托管栈,迁移期间两套框架并行运行,由代理 worker 按路由和用户分流,最终 Next.js 专属代码只占 3%。

    Awaiting translation

    Why it matters: Lovable 官方复盘把 90 万行代码从 Next.js 迁到自家 TanStack Start 栈,双框架并行、AI 智能体批量迁移和一次 OOM 事故的细节都可直接借鉴。

8/16Sun
  1. Thorsten Ball · Register Spill28

    Joy & Curiosity #95:作者用 Amp 在 orbs 中远程完成全部开发,不再使用本地开发环境

    作者称过去两周在 orbs 中借助 Amp 远程完成了全部开发工作,包括新的实验性 provider 后端、产品内 bug 上报功能、orbs 的磁盘与内存告警、启动日志、一个藏在官网里的 Comet Busters 类彩蛋游戏、听写功能的麦克风选择器、主题偏好等,并修复约二十个 bug、删除 5k 行代码。

    Awaiting translation

8/13Thu
  1. Augment Code · Blog24

    用 Cosmos 软件工厂打通工程与销售数据,看清企业试点真实健康度

    Augment Code 解决方案架构负责人用自家 AI 智能体编排平台 Cosmos 生成企业试点健康度视图,无需新建分析项目或仪表盘。Cosmos 会查询产品用量、核对会话血缘避免数字虚高、从 Salesforce 拉取交易状态,并对照客户通话记录验证结论,每个数字都可追溯到实时数据源。

    Awaiting translation

  2. Augment Code · Blog62

    Augment Code Expands Cosmos: Turning Code Review into an Agentic PR-to-Merge Loop

    Augment Code has extended its Cosmos review system from code review to a full PR-to-merge loop, adding four capabilities: Verifier, PR Fixer, Review Dashboard, and cosmos approve. Dedicated Experts handle risk analysis, line-by-line correctness review, design review, runtime verification, and fixes.

    Why it matters: Augment has expanded code review into a PR-to-merge loop covering fixes, verification, and approval, giving readers a way to judge how multi-agent division of labor plays out in practice.

8/12Wed
8/11Tue
  1. Lovable · Blog36

    Lovable 谈为什么"模型选择器"是死路:模型无关而非模型无差别

    Lovable 认为让用户在开工前选模型是死路,主张"模型无关":为每个模型定制指令、工具和项目上下文,而非套同一界面。其控制平面会按构建进展把不同环节分派给不同模型,并权衡切换带来的上下文损失。一次评测中,某前沿模型比前代任务完成快 15%、轮次少 40%、得分高 2–3%。

    Awaiting translation

8/9Sun
8/5Wed
  1. Vercel · v0 Blog62

    Vercel ships v0 API for programmatic access to its app-generation agent

    Vercel ships v0 API, giving programmatic, headless access to the v0 app-generation agent: send a prompt, v0 generates an app, spins up a dev server in the Vercel Sandbox, and returns a preview URL you can embed in your own UI. The API is now generally available.

    Why it matters: v0 opens up its app-generation capability as an API, so readers can judge how to wire it into their own product or agent workflow.

8/4Tue
8/3Mon
  1. OpenAI · Codex Cookbook71

    Iterating on a Development Workflow with Codex: From AGENTS.md to Phased Build Files

    The OpenAI Codex Cookbook lays out a set of repository conventions for wiring Codex into your development process: use AGENTS.md for persistent repository instructions, PLANS.md as the source of phase plans, and split the work into phased build files under harness/build/, with each phase spelling out its goals, acceptance criteria, boundaries, and approval gates.

    Why it matters: OpenAI lays out a complete directory convention for constraining Codex with AGENTS.md, PLANS.md, and phased build files—one you can adapt to your own repository.

8/2Sun
7/30Thu
  1. Terminal-Bench · News60

    Terminal-Bench 3.0 is out: 74 tasks across 7 domains, with the strongest model passing about 34%

    The Terminal-Bench team releases Terminal-Bench 3.0, whose first version spans 7 domains and 74 tasks, with the strongest model passing about 34%. Building on Terminal-Bench 2.1, this release broadens task diversity and adds CI/CD, semantic versioning, and result migration to keep improving the benchmark.

    Why it matters: Terminal-Bench 3.0 rebuilds the benchmark with 74 tasks and CI/CD-based versioning, so readers can see how the new benchmark separates models.

  2. Terminal-Bench · News62

    Terminal-Bench ships new Harbor features, turning the benchmark into a versioned asset that keeps getting updated

    The Terminal-Bench team has shipped a batch of new Harbor features that let datasets be released by version and let leaderboards migrate to new versions by reusing, re-evaluating, or rerunning trials. Tasks use semantic versioning: patch-level changes reuse old results as-is, validator changes only require re-evaluating saved artifacts, and only major changes that significantly alter the agent environment require a rerun. Dataset versions follow the highest version number among the tasks, and leaderboards use diffs to rerun only the tasks with major changes.

    Why it matters: The Terminal-Bench team maintains the benchmark like software, laying out concrete mechanisms for task semantic versioning and leaderboard upgrades that you can carry over to your own evaluation pipeline.

7/29Wed
  1. Lovable · Blog38

    Lovable 如何保障生产应用中的连接数据安全

    Lovable 为生产应用提供按用户授权的数据连接:每个人用自己的账号接入,只能看到源系统里本就有权查看的记录。凭据由服务端加密存储,应用不持有真实凭据,只向网关提交短期密钥和意图,网关再补上真实凭据转发。连接器被固定到唯一目标地址,请求无法把凭据发往错误服务器。

    Awaiting translation

7/25Sat
  1. Cline · Blog79

    Cline 用递归自我改进让智能体自主把 Kimi K3 在 Terminal-Bench 2.1 提到 88.8%

    Cline 用一条提示词启动 17 小时连续运行的编码智能体,以 GPT-5.6-Sol 为 leader 模型,把 Kimi K3 在 Terminal-Bench 2.1 上的成绩从基线 69/89(77.5%,$79)提升到确认运行 79/89(88.8%,$49.8),超过 Moonshot 自报的 88.3%。

    Awaiting translation

    Why it matters: Cline 用一条提示词让智能体自主完成 17 小时爬山,把 Terminal-Bench 2.1 从 77.5% 提到 88.8%,可看具体修了哪些 harness 问题。

7/24Fri
  1. Lovable · Blog71

    Lovable 如何用 AI 黑客智能体集群对自己打夺旗赛

    Lovable 在内部搭建了一套进攻性安全程序,让 AI 智能体集群像人类攻击者一样探测系统入口,直到拿到可验证的漏洞证据。它用夺旗赛的思路做验证:把 flag 散布在基础设施和权限最高的产品界面中,不预埋任何漏洞,智能体取到 flag 就说明找到了真实入侵路径,而不是模型猜测。

    Awaiting translation

    Why it matters: Lovable 公开了用夺旗机制验证漏洞的内部攻防智能体编排方法,可迁移到自家安全测试流程。

7/22Wed
  1. Lovable · Blog28

    Lovable 成为首个获得 AIUC-1 认证的 AI 编程智能体平台

    Lovable 成为首个获得 AIUC-1 认证的 AI 编程智能体平台,该标准是业界首个专为 AI 智能体打造的安全、保障与可靠性标准。AIUC-1 由 Stanford、MIT、MITRE 和云安全联盟参与制定,包含六大原则下的 51 项要求,覆盖密钥管理、安全代码生成默认设置、沙箱执行、人工监督和企业治理,每项要求均需提供证据并由第三方独立验证,而非自我声明。

    Awaiting translation

  2. Augment Code · Blog62

    What is loop engineering, and how are leading software engineering teams using it?

    Augment Code proposes loop engineering: designing agent loops that run from trigger to execution to validation to outcome, with agents handling the intermediate steps and humans stepping in only at checkpoints that require judgment. The article compares loop engineering with prompt engineering and context engineering as distinct layers, lays out five stages—trigger, execution, validation, outcome, and improvement—and describes four team-level loops already running in production: code review, ticket-to-PR, vulnerability remediation, and incident response.

    Why it matters: Augment Code breaks loop engineering into five stages—trigger, execution, validation, outcome, and improvement—and lays out four team-level loop patterns already running in production.

7/20Mon
  1. OpenAI Developer Blog · Codex71

    Codex Code Review now supports custom review rules in AGENTS.md

    OpenAI has added custom repository rules to Codex Code Review: you can put review guidelines in AGENTS.md, and Codex applies them during review and cites where each one came from in its findings. In OpenAI's own evaluation, the rule-guided version caught 98% of the required custom issues, versus 58.3% for the baseline. The guidance is to start with non-obvious invariants like compatibility requirements and data boundaries, put repo-level rules in the root directory and service-level rules in the corresponding directory, and leave formatting and mechanical checks to CI.

    Why it matters: OpenAI lays out the capabilities, the syntax, and the evaluation data for Codex Code Review custom rules, so you can judge how to bake your team's review experience into AGENTS.md.

7/18Sat
  1. Thorsten Ball · Register Spill22

    Joy & Curiosity #92:Amp 推出订阅与智能体间通信

    Amp 现已推出订阅服务,可与 ChatGPT 订阅搭配获得无限 GPT-5.6 token;同时 Amp 上线智能体间通信,智能体能在任意 Amp 实例或 orb 中派生其他智能体并互发消息与文件。作者还分享了新一季 Raising An Agent 播客、与 Evan Phoenix 等人的对谈,以及 antirez 关于“控制想法而非代码”的观点。

    Awaiting translation

7/15Wed
7/11Sat
7/8Wed