Skip to content
10/6 · Tue

Latest curated items

Today10/6Tue
  1. GitHub Blog · Copilot66

    GitHub 发布 AI 代码评审开放基准 ReviewBench

    GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。

    Awaiting translation

    Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。

10/5Mon
  1. AI Hero · Skills Updates62

    AI Hero Skills v1.3 released: new /implement-spec, /pr, and /retro skills, and CONTEXT.md renamed to GLOSSARY.md

    AI Hero Skills v1.3 is out, adding three skills—/implement-spec, /pr, and /retro—and extending the main flow from /grill-with-docs → /to-spec → /to-tickets into implementation, PRs, and retros.

    Why it matters: The author rounds out the skill set into a full path from writing specs to PRs and retros, and lays out the trade-offs at each step along with the rough edges they already know about.

10/4Sun
  1. 宝玉82

    Drawing on a podcast episode, Baoyu walks through how Lauren Tan, who works on Grok Bot at SpaceXAI, merged 2500 PRs in a single month: at night she lets the AI check and merge on its own, then spot-checks the next morning instead of reviewing each one.

    Quotedlauren@poteto

    i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful interviewer so i think the interview turned out really interesting both of our skill plugins work great together, so i recommend giving both a try and picking the best skills that suit your workflow https://www.youtube.com/watch?v=MN9dGgmLyso

    Why it matters: Using Lauren Tan's practice of merging 2500 PRs in a month, Baoyu explains that skipping individual reviews rests on validation Skills and rule constraints, and lays out the conditions under which he'd apply the same approach.

9/24Thu
  1. Lovable · Blog71

    Lovable ships Chats, and opens up the trajectory and inbox architecture behind its multi-agent collaboration

    Lovable launches Chats, an agent that runs at the workspace level, can hold conversations across projects and trigger builds; once changes are confirmed, it hands the task off to the project's builder agent and brings progress back into the conversation.

    Why it matters: Lovable shares the three-layer architecture behind Chats—trajectory, inbox, and activation—which you can adapt for your own multi-agent orchestration.

9/23Wed
  1. Cursor · Changelog62

    Cursor Launches Two Bots: Rollouts and Security Review

    Cursor has released two bots, Rollouts and Security Review, both available on Team and Enterprise plans.

    Why it matters: The official docs cover the monitoring and security review workflows for both bots, so readers can judge whether they fit into their existing delivery pipeline.

  2. Lovable · Blog60

    Lovable Ships Opus 5.5: Faster Builds, Quality on Par with Opus 5

    Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.

    Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.

9/15Tue
  1. Lovable · Blog64

    Lovable open-sources OJ, a Rust preview engine that beats Vite on cold start and memory

    Lovable has released OJ, a preview engine written from scratch in Rust. It reads your existing vite.config.ts and runs real Vite plugins through a compatibility layer, all in a single binary, with no toolchain installed into the project.

    Why it matters: Lovable rewrote its preview engine OJ in Rust, sharing cold start and memory comparisons against Vite, plus canary data from production.

  2. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/11Fri
  1. OpenAI Developer Blog · Codex66

    OpenAI on How to Rewrite Skills and Prompts for GPT-6 Astra

    In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.

    Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.

9/10Thu
  1. Cursor · Changelog76

    Cursor launches “Projects,” a feature that uses a coordinating agent to take on long-running development work

    Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.

    Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.

9/5Sat
  1. GitHub Blog · Copilot71

    GitHub Copilot launches Project HydraFusion, using multi-model runtime orchestration to improve coding quality

    GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.

    Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.

9/4Fri
  1. OpenAI Developer Blog · Codex71

    How to Build a Game with Astra in Codex: From Void Explorer to Performance Tuning

    The author built the space exploration game Void Explorer in Codex with Astra, featuring 2,048 star systems and over 10,000 procedurally generated planets, and shared the full workflow from prompts to architecture, testing, and performance measurement.

    Why it matters: Using Astra in Codex, the author built an entire game and showed a transferable collaborative workflow that spans prompts, testing, and performance measurement.

9/3Thu
  1. Cline · Blog74

    Cline 如何把 1100 万用户迁移到最大一次 harness 升级

    Cline 把 VS Code 扩展从约 76,000 行单体核心迁移到 Cline SDK,并自建灰度发布机制:一个安装包内打包 loader、legacy 和 next 两套扩展,由 PostHog 功能开关按百分比决定激活哪套,崩溃时自动回退到 legacy,开关可随时降到 0% 作为 kill switch。

    Awaiting translation

    Why it matters: Cline 官方复盘如何把 1100 万用户的 VS Code 扩展迁到新 harness,含灰度机制与前后指标对比。

  2. GitHub Blog · Copilot72

    GitHub Copilot 如何在不牺牲任务质量的前提下降低 AI 编码成本

    GitHub Copilot 团队复盘了四项降低 AI 编码成本的改动,核心原则是按完整任务而非单次工具调用衡量效率。

    Awaiting translation

    Why it matters: GitHub Copilot 团队复盘四项降本改动,并给出可迁移的评估方法:按完整任务而非单次工具调用衡量成本。

9/2Wed
  1. Cursor · Changelog66

    Cursor launches self-hosted machines, keeping tool execution within your own network

    Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.

    Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.

8/27Thu
  1. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。

8/26Wed
  1. Cline · Blog71

    Building a Code Review Agent on the Cline Loop with the Cline SDK

    The Cline team built a code review agent with the Cline SDK, splitting review into two agent loops—review and judge—then using a driver script to batch-submit the surviving issues as a single COMMENT event to the GitHub PR.

    Why it matters: A full breakdown of the plugin, Hooks, and two-stage loop behind a code review agent, transferable to other automated review scenarios.

8/25Tue
  1. Lovable · Blog62

    How Lovable connected its own app to external tech stacks: from MCP to nearly 100 connectors

    Lovable shared a retrospective on how it connected its platform app to third-party services: first it supported MCP as a stopgap for pulling context into chats, then it built app connectors of its own, using a Connector Gateway to proxy requests between published apps and third-party APIs. The gateway holds credentials and refresh logic, so deployed apps never touch the keys.

    Why it matters: Lovable’s retrospective on turning connectors into reusable infrastructure is worth a look for teams doing third-party integrations and credential management.

  2. OpenAI Developer Blog · Codex62

    Automating OpenAI’s repetitive evaluation work with Codex and the Runme notebook

    OpenAI engineers use Codex with the open-source notebook app Runme to automate repetitive work such as running model evaluations. The approach: write a goal cell in the Runme notebook, have Codex read the goal, produce a plan, and wait for human approval before executing, logging commands, outputs, and conclusions along the way—including the dead ends.

    Why it matters: The author uses the Runme notebook plus WebMCP to hand the evaluation process over to Codex; readers can borrow the way it handles goals, approvals, and context capture.

8/21Fri
  1. OpenAI Developer Blog · Codex65

    OpenAI Releases Daybreak and Codex Security, a Security Workflow

    OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.

    Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.