OpenAI 在欧盟为 ChatGPT 和 Codex 文本加入 textGrain 隐藏水印
OpenAI 表示欧盟符合条件的 ChatGPT 和 Codex 文本将携带名为 textGrain 的隐藏水印,未来几周内面向所有套餐推出,目前仅限欧盟。该水印是词选择上的统计模式而非可见标签,OpenAI 称其不识别用户、账号或提示词,检测器不公开,仅获批研究人员可申请使用。
Awaiting translation
OpenAI 表示欧盟符合条件的 ChatGPT 和 Codex 文本将携带名为 textGrain 的隐藏水印,未来几周内面向所有套餐推出,目前仅限欧盟。该水印是词选择上的统计模式而非可见标签,OpenAI 称其不识别用户、账号或提示词,检测器不公开,仅获批研究人员可申请使用。
Awaiting translation
Future AGI 1.47.0 发布,在运行分析中新增通话与语音指标卡片,并修复模拟通话结果与运行详情的对齐问题。评估对话时,错误语言回答、涉及其他产品的回答以及未获回复的请求现统一计为未处理请求;测试环境评估还修复了完整提示词和对话参与者标识的传递。聊天构建器页面支持折叠与调整宽度,宽度在窗口缩放后保留。
Awaiting translation
SemiAnalysis 实测 Anthropic、OpenAI 等九家 AI 订阅套餐后得出,同样 200 美元,Claude 订阅折算的 Token 用量约为 OpenAI 的 5 倍。
Awaiting translation
Anthropic Subscriptions Offer 5x+ More Value Than OpenAI Limit testing every AI subscription plan from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Zdotai, Cursor, and Cognition https://newsletter.semianalysis.com/p/anthropic-subscriptions-offer-5x
OpenAI 在 9 月 22 日至 10 月 1 日间连续发布 Codex CLI 0.156.0 到 0.160.0,官方称之为一次大更新,带来全屏终端界面、/agents 视图和并行任务管理。
Awaiting translation
GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。
Awaiting translation
Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。
微软与加州大学圣巴巴拉分校研究者公开论文 ScholarEvolve,从已发表的 Agent 研究中寻找改进思路,写进 Harness 后用真实任务检验,执行任务的模型保持不变。
Awaiting translation
GPT-5.5 于 2026 年 4 月 23 日发布,是 OpenAI 自 GPT-4.5 以来首个完全重训练的基座模型,在 Terminal-Bench 2.0 上得分 82.7%,相同 Codex 任务下比 GPT-5.4 少用 40% token,但输入输出价格翻倍至每百万 token 5 美元和 30 美元。
Awaiting translation
Cursor 上周扩展 Cloud Agents,新增自托管机器、团队 worker 池和可长期运行的 Projects。
Awaiting translation
Cloudflare 发布 Web Search API,通过 AI Gateway 让 AI 智能体的回答接入实时数据,而非依赖训练数据猜测。测试版首批支持 Ceramic.ai、Exa 和 Linkup 三家搜索提供商,均提供 Zero Data Retention 模式。
Awaiting translation
作者上手体验了 OpenAI DevDay 发布的 Dots、Decisions API、开源 Codex Harness 更新和 ChatGPT Space 四项能力。
Awaiting translation
On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.
Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.
JetBrains 为 IDE 内的 Air 开启早期访问计划,插件已上架 Marketplace,2026.3 EAP 版本中已内置。
Awaiting translation
Google DeepMind 发布 Gemini 4 Argon,单次输出上限从 64K 提升到 1M tokens,主打长任务与多步骤推理。
Awaiting translation
Why it matters: 汇总了 Gemini 4 Argon 的官方评测数据与 Google 内部落地案例,可对照各家旗舰模型的能力差异。
谷歌发布旗舰模型 Gemini 4 Argon,官方公布的 18 项测试中独占第一 12 项,DeepSWE v1.1 以 77.9% 位列全球第一,AutomationBench-AA 以 77.5% 领先 Claude Sonnet 5.5 的 71.3%。
Awaiting translation
OpenAI 在 9 月 29 日 DevDay 上发布 GPT-6.1 Sol,距离 9 月 22 日发布 GPT-6 Sol 仅一周,官方测试中接近 Astra 水准,API 标准输入输出单价为 Astra 的 1/5。
Awaiting translation
Simon Willison live-blogged the OpenAI DevDay 2026 keynote from Fort Mason in San Francisco, where OpenAI announced the personal agent Dots, ChatGPT Space, GPT-6.1 Sol, Ultrafast, and more.
Why it matters: A running, item-by-item record of what OpenAI announced at DevDay, for a quick look at what Dots, GPT-6.1 Sol, Ultrafast, and Codex Security actually look like.
OpenAI 在 9 月 29 日旧金山 DevDay 上发布 GPT-6.1 Sol,API 名为 gpt-6.1-sol,定价为每百万输入 token 2 美元、输出 10 美元,缓存输入 0.10 美元,标准价格是 GPT-6 Astra 的五分之一。
Awaiting translation
Why it matters: OpenAI DevDay 发布 GPT-6.1 Sol,价格降至 Astra 的五分之一,并同步更新 Codex、Agents API 与插件体系,可据此判断成本与工具链变化。
Cloudflare 发布 Worker Previews,让每个 Git 分支拥有独立且接近生产的环境,各自有稳定 URL、独立配置和状态,并带独立可观测性。
Awaiting translation
n8n 2.42.0 预览版于 9 月 29 日发布,默认启用智能体,并为智能体新增共享消息队列与共享上下文访问,预览和集成也接入该队列。此次更新影响 Agent Builder 用户和集成开发者,编辑器新增记录锁定、预览队列消息管理和渠道内智能体操作确认设置;OAuth2 新增 n8n User Auth 的 Webhook 浏览器授权流程。
Awaiting translation
OpenAI 发布 Codex CLI 0.157.0,主要变化是工具会为支持的交互式会话自动启动后台服务器,让 Codex 的使用不再绑定在单个终端窗口。新增的 f 键可以从 CLI 分叉其他应用中打开的对话,保留草稿和排队中的提示词,/import 命令也能在远程和本地后台会话中使用。原文指出这并不代表关闭终端后任务仍会继续执行,OpenAI 没有这样的声明,后台服务器只是为共享会话打基础。
Awaiting translation
Awaiting translation
Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels
Visual Studio Code 1.139 稳定版发布,AI 智能体现在可以在 SSH、Tunnel 和 WSL 远程主机的 Dev Containers 中构建和测试项目,使用容器内配置的工具与依赖。
Awaiting translation
GitHub 于 9 月 23 日为 GitHub Copilot App 推出本地沙箱,目前处于公开预览,仅适用于使用本地仓库和工作树的会话。沙箱可分别限制文件系统读写路径、外网与局域网访问,以及 Git credentials 和 GitHub CLI 数据的使用,企业策略还能进一步收紧;若操作系统无法应用规则,隔离环境会报错退出而不执行命令。
Awaiting translation
On September 22, Anthropic released its flagship model Claude Opus 5.5, aimed at developers and teams who want agents to handle multi-step tasks like coding and data analysis. The company says it delivers better performance and lower cost than Opus 5.
Why it matters: Anthropic's published pricing and the default workload cost reduction help developers estimate the migration cost for long-running agent tasks.
Cursor has released two bots, Rollouts and Security Review, both available on Team and Enterprise plans.
Why it matters: The official docs cover the monitoring and security review workflows for both bots, so readers can judge whether they fit into their existing delivery pipeline.
Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.
Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.
Awaiting translation
MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai
Claude Code 从 2.1.277 版本开始支持 AGENTS.md:当文件夹中没有 CLAUDE.md 时,Claude 会检查并使用 AGENTS.md。该支持基于 Claude Code mods 构建,这是其即将推出的定制 Claude Code harness 的方式,属于内置 mod,用户之后也可以自行构建自定义版本的项目指令。
Awaiting translation
Lovable has released OJ, a preview engine written from scratch in Rust. It reads your existing vite.config.ts and runs real Vite plugins through a compatibility layer, all in a single binary, with no toolchain installed into the project.
Why it matters: Lovable rewrote its preview engine OJ in Rust, sharing cold start and memory comparisons against Vite, plus canary data from production.
Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.
Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.
开发者 pdfu 在 iOS 27 和 macOS Golden Gate 的私有框架中发现,Apple 的新 Siri 架构设计了与第三方 AI 模型对接的机制。
Awaiting translation
Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.
Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.
GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.
Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.
On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.
Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.
Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.
Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.
OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.
Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.
Simon Willison 用 Claude Code for web 做了一个约 150 行、零依赖的 TypeScript 服务原型,基于 Bun 1.4 实验性的 Bun.WebView 提供 shot-scraper 风格的 JSON API,支持执行 JavaScript 和输出 PNG/JPEG/WebP 截图,无需 Puppeteer 或 Playwright。
Awaiting translation
Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.
Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.
OpenAI has open-sourced the harness that drives the Codex app, CLI, and IDE extensions, and through the Codex app-server client protocol it exposes capabilities like creating threads, starting turns, receiving events, and handling approval requests.
Why it matters: With the Codex harness and app-server protocol now public, developers can see how to embed the agent in their own products and where the boundaries are.
NVIDIA has released Nemotron 3.5 Lightning, a customizable open-source model built for persistent agents, and it's now available for free in Cline.
Why it matters: NVIDIA's new open-source model is free to use on Cline, so you can decide whether it's worth switching for high-frequency agent workloads.