Skip to content

#Agent

0 items today
10/6Tue
  1. 宝玉71

    SemiAnalysis 实测 Anthropic、OpenAI 等九家 AI 订阅套餐后得出,同样 200 美元,Claude 订阅折算的 Token 用量约为 OpenAI 的 5 倍。

    Awaiting translation

    QuotedSemiAnalysis@SemiAnalysis_

    Anthropic Subscriptions Offer 5x+ More Value Than OpenAI Limit testing every AI subscription plan from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Zdotai, Cursor, and Cognition https://newsletter.semianalysis.com/p/anthropic-subscriptions-offer-5x

  2. 宝玉71

    据 The Information 10 月 5 日报道,Meta 和微软都在减少员工内部使用 Anthropic 的 Claude,转向自家模型和工具。

    Awaiting translation

    QuotedNIK@ns123abc

    🚨BREAKING: Microsoft and META are aggressively cutting employee use of Claude ahead of Anthropic's IPO Microsoft has cut internal claude spend by more than 33%, nuked per-employee token budget from $100k/month to $10k/month, and forced Copilot to auto-route to cheaper models META used Claude code to build Muse, then cut active users from 60,000 to 30,000 (50% decline) after launch, and replaced it with Muse Code Palantir and Nvidia are also scaling back claude over soaring prices and data privacy fears it’s OVER…

  3. 宝玉70

    amontlabs/lcu 把 Codex 的 Computer Use 单独拆出,让 Claude Code、Codex CLI、Pi 等 Agent 工具通过 MCP 调用。

    Awaiting translation

    Quoted向阳乔木@vista8

    发现一个牛逼的东西,让任意 Agent 调用 Codex 的 Computer Use。 Codex 最强的就是 Computer Use。 但最近用 Claude Opus 5.5比较多,这样就强强联合了。 刚测试通过,安装后建议配置个 Hook,指定白名单可控制哪些 App 安装地址见评论区

  4. GitHub Blog · Copilot66

    GitHub 发布 AI 代码评审开放基准 ReviewBench

    GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。

    Awaiting translation

    Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。

10/4Sun
  1. 宝玉82

    Drawing on a podcast episode, Baoyu walks through how Lauren Tan, who works on Grok Bot at SpaceXAI, merged 2500 PRs in a single month: at night she lets the AI check and merge on its own, then spot-checks the next morning instead of reviewing each one.

    Quotedlauren@poteto

    i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful interviewer so i think the interview turned out really interesting both of our skill plugins work great together, so i recommend giving both a try and picking the best skills that suit your workflow https://www.youtube.com/watch?v=MN9dGgmLyso

    Why it matters: Using Lauren Tan's practice of merging 2500 PRs in a month, Baoyu explains that skipping individual reviews rests on validation Skills and rule constraints, and lays out the conditions under which he'd apply the same approach.

10/3Sat
  1. 宝玉62

    剑桥大学 AI 科学与政策项目(CASP)9 月发布一篇论文,Hinton、Bengio 等 20 多位研究者联名讨论 AI 研发被自动化后是否会引发智能爆炸,Hinton 在 X 上推荐了该论文。

    Awaiting translation

    QuotedGeoffrey Hinton@geoffreyhinton

    The idea of an intelligence explosion caused by recursive self improvement has been around for a long time but until very recently it did not seem imminent. Now many leading researchers think it may happen quite soon. You can read our paper about it here: https://casp.ac/reports/intelligence-explosion

10/1Thu
9/26Sat
  1. Boris Cherny60

    Claude Tag 在 Slack 中现已支持个人连接器,可直接访问个人有权限的 Drive 文档、Salesforce 账号或数仓表,今天在 Teams 上线、下周面向 Enterprise 开放。

    Awaiting translation

    QuotedNoah Zweben@noahzweben

    Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels

9/25Fri
  1. GitHub Blog · Copilot34

    GitHub Copilot 的 canvas:当聊天框不是与 AI 协作的正确界面

    GitHub Copilot 应用推出 canvas,一种运行在应用内、无浏览器外壳的全栈小应用,可与 Copilot 智能体双向通信,也能在本地执行代码。作者认为 chat 只适合作为通用兜底界面,用户明确任务时更该让智能体生成可复用工具,例如用 canvas 管理 Winget 包或直接操作 SQLite 数据库,避免把 token 浪费在“stage and commit”这类操作上。

    Awaiting translation

9/24Thu
  1. Lovable · Blog71

    Lovable ships Chats, and opens up the trajectory and inbox architecture behind its multi-agent collaboration

    Lovable launches Chats, an agent that runs at the workspace level, can hold conversations across projects and trigger builds; once changes are confirmed, it hands the task off to the project's builder agent and brings progress back into the conversation.

    Why it matters: Lovable shares the three-layer architecture behind Chats—trajectory, inbox, and activation—which you can adapt for your own multi-agent orchestration.

9/23Wed
  1. Lovable · Blog60

    Lovable Ships Opus 5.5: Faster Builds, Quality on Par with Opus 5

    Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.

    Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.

9/22Tue
  1. Sebastian Raschka72

    小米发布开源权重模型 MiMo-V2.6-Pro,在 Artificial Analysis 智能指数上以 46 分成为开源权重模型第一,每任务成本 0.13 美元,输入 0.435 美元/1M tokens、输出 0.87 美元/1M tokens,采用 1.02T 总参数、42B 激活参数的 MoE 架构。

    Awaiting translation

    QuotedArtificial Analysis@ArtificialAnlys

    MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://artificialanalysis.ai

9/15Tue
  1. Lovable · Blog64

    Lovable open-sources OJ, a Rust preview engine that beats Vite on cold start and memory

    Lovable has released OJ, a preview engine written from scratch in Rust. It reads your existing vite.config.ts and runs real Vite plugins through a compatibility layer, all in a single binary, with no toolchain installed into the project.

    Why it matters: Lovable rewrote its preview engine OJ in Rust, sharing cold start and memory comparisons against Vite, plus canary data from production.

  2. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/12Sat
9/11Fri
  1. OpenAI Developer Blog · Codex66

    OpenAI on How to Rewrite Skills and Prompts for GPT-6 Astra

    In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.

    Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.

9/10Thu
  1. Cursor · Changelog76

    Cursor launches “Projects,” a feature that uses a coordinating agent to take on long-running development work

    Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.

    Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.

9/5Sat
  1. GitHub Blog · Copilot71

    GitHub Copilot launches Project HydraFusion, using multi-model runtime orchestration to improve coding quality

    GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.

    Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.

9/4Fri
  1. OpenAI Developer Blog · Codex71

    How to Build a Game with Astra in Codex: From Void Explorer to Performance Tuning

    The author built the space exploration game Void Explorer in Codex with Astra, featuring 2,048 star systems and over 10,000 procedurally generated planets, and shared the full workflow from prompts to architecture, testing, and performance measurement.

    Why it matters: Using Astra in Codex, the author built an entire game and showed a transferable collaborative workflow that spans prompts, testing, and performance measurement.

9/3Thu
  1. Cline · Blog74

    Cline 如何把 1100 万用户迁移到最大一次 harness 升级

    Cline 把 VS Code 扩展从约 76,000 行单体核心迁移到 Cline SDK,并自建灰度发布机制:一个安装包内打包 loader、legacy 和 next 两套扩展,由 PostHog 功能开关按百分比决定激活哪套,崩溃时自动回退到 legacy,开关可随时降到 0% 作为 kill switch。

    Awaiting translation

    Why it matters: Cline 官方复盘如何把 1100 万用户的 VS Code 扩展迁到新 harness,含灰度机制与前后指标对比。

9/2Wed
  1. Cursor · Changelog66

    Cursor launches self-hosted machines, keeping tool execution within your own network

    Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.

    Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.

8/30Sun
8/26Wed
  1. Cline · Blog71

    Building a Code Review Agent on the Cline Loop with the Cline SDK

    The Cline team built a code review agent with the Cline SDK, splitting review into two agent loops—review and judge—then using a driver script to batch-submit the surviving issues as a single COMMENT event to the GitHub PR.

    Why it matters: A full breakdown of the plugin, Hooks, and two-stage loop behind a code review agent, transferable to other automated review scenarios.

8/25Tue
  1. Lovable · Blog62

    How Lovable connected its own app to external tech stacks: from MCP to nearly 100 connectors

    Lovable shared a retrospective on how it connected its platform app to third-party services: first it supported MCP as a stopgap for pulling context into chats, then it built app connectors of its own, using a Connector Gateway to proxy requests between published apps and third-party APIs. The gateway holds credentials and refresh logic, so deployed apps never touch the keys.

    Why it matters: Lovable’s retrospective on turning connectors into reusable infrastructure is worth a look for teams doing third-party integrations and credential management.

  2. OpenAI Developer Blog · Codex62

    Automating OpenAI’s repetitive evaluation work with Codex and the Runme notebook

    OpenAI engineers use Codex with the open-source notebook app Runme to automate repetitive work such as running model evaluations. The approach: write a goal cell in the Runme notebook, have Codex read the goal, produce a plan, and wait for human approval before executing, logging commands, outputs, and conclusions along the way—including the dead ends.

    Why it matters: The author uses the Runme notebook plus WebMCP to hand the evaluation process over to Codex; readers can borrow the way it handles goals, approvals, and context capture.

8/21Fri
  1. OpenAI Developer Blog · Codex65

    OpenAI Releases Daybreak and Codex Security, a Security Workflow

    OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.

    Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.

8/19Wed
  1. Cursor · Changelog66

    Cursor Updates Cloud Agents and Harness: Event Subscriptions, Independent Sub-Agent Runs, and /goal

    Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.

    Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.

  2. Cline · Blog71

    Cline's Evaluation Methodology and Trace for Open-Weight Models

    Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.

    Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.

  3. OpenAI Developer Blog · Codex71

    OpenAI open-sources the Codex harness and the app-server client protocol

    OpenAI has open-sourced the harness that drives the Codex app, CLI, and IDE extensions, and through the Codex app-server client protocol it exposes capabilities like creating threads, starting turns, receiving events, and handling approval requests.

    Why it matters: With the Codex harness and app-server protocol now public, developers can see how to embed the agent in their own products and where the boundaries are.

8/18Tue
  1. Lovable · Blog74

    Lovable 如何把 lovable.dev 从 Next.js 迁到自家 TanStack Start 栈

    Lovable 用六个月把月访问 4200 万的 lovable.dev 从 Next.js 迁到自家 TanStack Start 托管栈,迁移期间两套框架并行运行,由代理 worker 按路由和用户分流,最终 Next.js 专属代码只占 3%。

    Awaiting translation

    Why it matters: Lovable 官方复盘把 90 万行代码从 Next.js 迁到自家 TanStack Start 栈,双框架并行、AI 智能体批量迁移和一次 OOM 事故的细节都可直接借鉴。

8/13Thu
  1. Augment Code · Blog62

    Augment Code Expands Cosmos: Turning Code Review into an Agentic PR-to-Merge Loop

    Augment Code has extended its Cosmos review system from code review to a full PR-to-merge loop, adding four capabilities: Verifier, PR Fixer, Review Dashboard, and cosmos approve. Dedicated Experts handle risk analysis, line-by-line correctness review, design review, runtime verification, and fixes.

    Why it matters: Augment has expanded code review into a PR-to-merge loop covering fixes, verification, and approval, giving readers a way to judge how multi-agent division of labor plays out in practice.

8/11Tue
8/9Sun
8/5Wed
  1. Vercel · v0 Blog62

    Vercel ships v0 API for programmatic access to its app-generation agent

    Vercel ships v0 API, giving programmatic, headless access to the v0 app-generation agent: send a prompt, v0 generates an app, spins up a dev server in the Vercel Sandbox, and returns a preview URL you can embed in your own UI. The API is now generally available.

    Why it matters: v0 opens up its app-generation capability as an API, so readers can judge how to wire it into their own product or agent workflow.

8/3Mon
  1. OpenAI · Codex Cookbook71

    Iterating on a Development Workflow with Codex: From AGENTS.md to Phased Build Files

    The OpenAI Codex Cookbook lays out a set of repository conventions for wiring Codex into your development process: use AGENTS.md for persistent repository instructions, PLANS.md as the source of phase plans, and split the work into phased build files under harness/build/, with each phase spelling out its goals, acceptance criteria, boundaries, and approval gates.

    Why it matters: OpenAI lays out a complete directory convention for constraining Codex with AGENTS.md, PLANS.md, and phased build files—one you can adapt to your own repository.

8/2Sun
7/30Thu
  1. Terminal-Bench · News60

    Terminal-Bench 3.0 is out: 74 tasks across 7 domains, with the strongest model passing about 34%

    The Terminal-Bench team releases Terminal-Bench 3.0, whose first version spans 7 domains and 74 tasks, with the strongest model passing about 34%. Building on Terminal-Bench 2.1, this release broadens task diversity and adds CI/CD, semantic versioning, and result migration to keep improving the benchmark.

    Why it matters: Terminal-Bench 3.0 rebuilds the benchmark with 74 tasks and CI/CD-based versioning, so readers can see how the new benchmark separates models.