Skip to content

#Agent

0 items today
10/6Tue
  1. Reddit · ClaudeCode / Codex / VibeCoding60

    OpenAI 在欧盟为 ChatGPT 和 Codex 文本加入 textGrain 隐藏水印

    OpenAI 表示欧盟符合条件的 ChatGPT 和 Codex 文本将携带名为 textGrain 的隐藏水印,未来几周内面向所有套餐推出,目前仅限欧盟。该水印是词选择上的统计模式而非可见标签,OpenAI 称其不识别用户、账号或提示词,检测器不公开,仅获批研究人员可申请使用。

    Awaiting translation

  2. Tproger · Программирование12

    Future AGI 1.47.0 新增通话指标并改进 AI 智能体回答评估

    Future AGI 1.47.0 发布,在运行分析中新增通话与语音指标卡片,并修复模拟通话结果与运行详情的对齐问题。评估对话时,错误语言回答、涉及其他产品的回答以及未获回复的请求现统一计为未处理请求;测试环境评估还修复了完整提示词和对话参与者标识的传递。聊天构建器页面支持折叠与调整宽度,宽度在窗口缩放后保留。

    Awaiting translation

  3. 宝玉71

    SemiAnalysis 实测 Anthropic、OpenAI 等九家 AI 订阅套餐后得出,同样 200 美元,Claude 订阅折算的 Token 用量约为 OpenAI 的 5 倍。

    Awaiting translation

    QuotedSemiAnalysis@SemiAnalysis_

    Anthropic Subscriptions Offer 5x+ More Value Than OpenAI Limit testing every AI subscription plan from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Zdotai, Cursor, and Cognition https://newsletter.semianalysis.com/p/anthropic-subscriptions-offer-5x

  4. GitHub Blog · Copilot66

    GitHub 发布 AI 代码评审开放基准 ReviewBench

    GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。

    Awaiting translation

    Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。

10/5Mon
10/4Sun
10/3Sat
  1. TonyBai80

    Pi 1.0 is out: the Agent engine behind OpenClaw now takes in MCP and ships Pi Durable with crash recovery

    On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.

    Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.

10/2Fri
10/1Thu
9/30Wed
9/29Tue
  1. Simon Willison · Coding Agents78

    Live from OpenAI DevDay 2026: Dots personal agent, GPT-6.1 Sol, and Ultrafast announced

    Simon Willison live-blogged the OpenAI DevDay 2026 keynote from Fort Mason in San Francisco, where OpenAI announced the personal agent Dots, ChatGPT Space, GPT-6.1 Sol, Ultrafast, and more.

    Why it matters: A running, item-by-item record of what OpenAI announced at DevDay, for a quick look at what Dots, GPT-6.1 Sol, Ultrafast, and Codex Security actually look like.

  2. Tproger · Программирование88

    OpenAI 在 DevDay 发布 GPT-6.1 Sol,价格仅为 Astra 的五分之一

    OpenAI 在 9 月 29 日旧金山 DevDay 上发布 GPT-6.1 Sol,API 名为 gpt-6.1-sol,定价为每百万输入 token 2 美元、输出 10 美元,缓存输入 0.10 美元,标准价格是 GPT-6 Astra 的五分之一。

    Awaiting translation

    Why it matters: OpenAI DevDay 发布 GPT-6.1 Sol,价格降至 Astra 的五分之一,并同步更新 Codex、Agents API 与插件体系,可据此判断成本与工具链变化。

  3. Tproger · Программирование22

    n8n 2.42.0 预览版默认启用智能体

    n8n 2.42.0 预览版于 9 月 29 日发布,默认启用智能体,并为智能体新增共享消息队列与共享上下文访问,预览和集成也接入该队列。此次更新影响 Agent Builder 用户和集成开发者,编辑器新增记录锁定、预览队列消息管理和渠道内智能体操作确认设置;OAuth2 新增 n8n User Auth 的 Webhook 浏览器授权流程。

    Awaiting translation

9/26Sat
  1. Habr · Codex58

    OpenAI 发布 Codex CLI 0.157.0,自动启动后台服务

    OpenAI 发布 Codex CLI 0.157.0,主要变化是工具会为支持的交互式会话自动启动后台服务器,让 Codex 的使用不再绑定在单个终端窗口。新增的 f 键可以从 CLI 分叉其他应用中打开的对话,保留草稿和排队中的提示词,/import 命令也能在远程和本地后台会话中使用。原文指出这并不代表关闭终端后任务仍会继续执行,OpenAI 没有这样的声明,后台服务器只是为共享会话打基础。

    Awaiting translation

  2. Boris Cherny60

    Claude Tag 在 Slack 中现已支持个人连接器,可直接访问个人有权限的 Drive 文档、Salesforce 账号或数仓表,今天在 Teams 上线、下周面向 Enterprise 开放。

    Awaiting translation

    QuotedNoah Zweben@noahzweben

    Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels

9/24Thu
  1. Tproger · Программирование58

    GitHub Copilot App 推出本地沙箱,限制文件、网络与凭据访问

    GitHub 于 9 月 23 日为 GitHub Copilot App 推出本地沙箱,目前处于公开预览,仅适用于使用本地仓库和工作树的会话。沙箱可分别限制文件系统读写路径、外网与局域网访问,以及 Git credentials 和 GitHub CLI 数据的使用,企业策略还能进一步收紧;若操作系统无法应用规则,隔离环境会报错退出而不执行命令。

    Awaiting translation

9/23Wed
  1. Tproger · Программирование80

    Anthropic Releases Flagship Model Claude Opus 5.5

    On September 22, Anthropic released its flagship model Claude Opus 5.5, aimed at developers and teams who want agents to handle multi-step tasks like coding and data analysis. The company says it delivers better performance and lower cost than Opus 5.

    Why it matters: Anthropic's published pricing and the default workload cost reduction help developers estimate the migration cost for long-running agent tasks.

  2. Lovable · Blog60

    Lovable Ships Opus 5.5: Faster Builds, Quality on Par with Opus 5

    Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.

    Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.

9/22Tue
  1. Sebastian Raschka72

    小米发布开源权重模型 MiMo-V2.6-Pro,在 Artificial Analysis 智能指数上以 46 分成为开源权重模型第一,每任务成本 0.13 美元,输入 0.435 美元/1M tokens、输出 0.87 美元/1M tokens,采用 1.02T 总参数、42B 激活参数的 MoE 架构。

    Awaiting translation

    QuotedArtificial Analysis@ArtificialAnlys

    MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai

9/19Sat
  1. Simon Willison · Coding Agents65

    Claude Code 2.1.277 起支持 AGENTS.md

    Claude Code 从 2.1.277 版本开始支持 AGENTS.md:当文件夹中没有 CLAUDE.md 时,Claude 会检查并使用 AGENTS.md。该支持基于 Claude Code mods 构建,这是其即将推出的定制 Claude Code harness 的方式,属于内置 mod,用户之后也可以自行构建自定义版本的项目指令。

    Awaiting translation

9/15Tue
  1. Lovable · Blog64

    Lovable open-sources OJ, a Rust preview engine that beats Vite on cold start and memory

    Lovable has released OJ, a preview engine written from scratch in Rust. It reads your existing vite.config.ts and runs real Vite plugins through a compatibility layer, all in a single binary, with no toolchain installed into the project.

    Why it matters: Lovable rewrote its preview engine OJ in Rust, sharing cold start and memory comparisons against Vite, plus canary data from production.

  2. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/14Mon
9/10Thu
  1. Cursor · Changelog76

    Cursor launches “Projects,” a feature that uses a coordinating agent to take on long-running development work

    Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.

    Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.

9/5Sat
  1. GitHub Blog · Copilot71

    GitHub Copilot launches Project HydraFusion, using multi-model runtime orchestration to improve coding quality

    GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.

    Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.

9/4Fri
  1. DevAgentStack · Field Notes82

    GPT-6 Astra vs. Fable 5.1 benchmarks: which scores are comparable and which aren't

    On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.

    Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.

9/2Wed
  1. Cursor · Changelog66

    Cursor launches self-hosted machines, keeping tool execution within your own network

    Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.

    Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.

8/21Fri
  1. OpenAI Developer Blog · Codex65

    OpenAI Releases Daybreak and Codex Security, a Security Workflow

    OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.

    Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.

8/20Thu
8/19Wed
  1. Cursor · Changelog66

    Cursor Updates Cloud Agents and Harness: Event Subscriptions, Independent Sub-Agent Runs, and /goal

    Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.

    Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.

  2. OpenAI Developer Blog · Codex71

    OpenAI open-sources the Codex harness and the app-server client protocol

    OpenAI has open-sourced the harness that drives the Codex app, CLI, and IDE extensions, and through the Codex app-server client protocol it exposes capabilities like creating threads, starting turns, receiving events, and handling approval requests.

    Why it matters: With the Codex harness and app-server protocol now public, developers can see how to embed the agent in their own products and where the boundaries are.

8/11Tue