Skip to content

Tool updates

Version releases and feature updates for AI coding tools.

Latest curated items

Items 1–20 · 30 total
10/5Mon
  1. AI Hero · Skills Updates62

    AI Hero Skills v1.3 released: new /implement-spec, /pr, and /retro skills, and CONTEXT.md renamed to GLOSSARY.md

    AI Hero Skills v1.3 is out, adding three skills—/implement-spec, /pr, and /retro—and extending the main flow from /grill-with-docs → /to-spec → /to-tickets into implementation, PRs, and retros.

    Why it matters: The author rounds out the skill set into a full path from writing specs to PRs and retros, and lays out the trade-offs at each step along with the rough edges they already know about.

10/3Sat
  1. TonyBai80

    Pi 1.0 is out: the Agent engine behind OpenClaw now takes in MCP and ships Pi Durable with crash recovery

    On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.

    Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.

9/28Mon
  1. GitHubDaily78

    Paseo: one interface to manage Claude Code, Codex, and other AI agents

    The open-source project Paseo pulls command-line agents like Claude Code, Codex, OpenCode, and Pi into a single management interface. All agents still run locally on your own machine, and the project has already reached 18700+ stars.

    Why it matters: The original article shows how to bring multiple command-line agents into one interface and hand off tasks across agents, making it a useful reference for developers running several agents at once.

9/24Thu
  1. Lovable · Blog71

    Lovable ships Chats, and opens up the trajectory and inbox architecture behind its multi-agent collaboration

    Lovable launches Chats, an agent that runs at the workspace level, can hold conversations across projects and trigger builds; once changes are confirmed, it hands the task off to the project's builder agent and brings progress back into the conversation.

    Why it matters: Lovable shares the three-layer architecture behind Chats—trajectory, inbox, and activation—which you can adapt for your own multi-agent orchestration.

9/23Wed
  1. Lovable · Blog60

    Lovable Ships Opus 5.5: Faster Builds, Quality on Par with Opus 5

    Lovable has shipped Opus 5.5, which the company says matches Opus 5 in results while cutting the number of steps by one-third to one-half. On Lovable's internal benchmarks, Opus 5.5 ties Opus 5 on 0-to-1 builds and iterative code changes, and comes out 4% to 6% ahead on validation discipline; across all reasoning effort levels, steps per task drop by 26% to 57% and input tokens fall by 21% to 59%, with the differences significant at the 95% confidence level.

    Why it matters: Lovable shares official comparison data between Opus 5.5 and Opus 5 on step counts and tokens, so readers can judge the real change in build efficiency.

9/15Tue
  1. Lovable · Blog64

    Lovable open-sources OJ, a Rust preview engine that beats Vite on cold start and memory

    Lovable has released OJ, a preview engine written from scratch in Rust. It reads your existing vite.config.ts and runs real Vite plugins through a compatibility layer, all in a single binary, with no toolchain installed into the project.

    Why it matters: Lovable rewrote its preview engine OJ in Rust, sharing cold start and memory comparisons against Vite, plus canary data from production.

  2. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/10Thu
  1. Cursor · Changelog76

    Cursor launches “Projects,” a feature that uses a coordinating agent to take on long-running development work

    Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.

    Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.

9/5Sat
  1. GitHub Blog · Copilot71

    GitHub Copilot launches Project HydraFusion, using multi-model runtime orchestration to improve coding quality

    GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.

    Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.

9/2Wed
  1. Cursor · Changelog66

    Cursor launches self-hosted machines, keeping tool execution within your own network

    Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.

    Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.

8/21Fri
  1. OpenAI Developer Blog · Codex65

    OpenAI Releases Daybreak and Codex Security, a Security Workflow

    OpenAI has launched Daybreak, combining ChatGPT, Codex Security, and the open-source Codex Security CLI into a security defense workflow that covers pre-merge PR reviews, repository and vulnerability backlog scans, and regular CI checks.

    Why it matters: The official documentation walks through the full Codex Security workflow—from PR reviews and repository scans to CLI-based batch scanning—so you can decide how to plug it into your existing security processes.

8/19Wed
  1. Cursor · Changelog66

    Cursor Updates Cloud Agents and Harness: Event Subscriptions, Independent Sub-Agent Runs, and /goal

    Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.

    Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.

  2. OpenAI Developer Blog · Codex71

    OpenAI open-sources the Codex harness and the app-server client protocol

    OpenAI has open-sourced the harness that drives the Codex app, CLI, and IDE extensions, and through the Codex app-server client protocol it exposes capabilities like creating threads, starting turns, receiving events, and handling approval requests.

    Why it matters: With the Codex harness and app-server protocol now public, developers can see how to embed the agent in their own products and where the boundaries are.

8/13Thu
  1. Augment Code · Blog62

    Augment Code Expands Cosmos: Turning Code Review into an Agentic PR-to-Merge Loop

    Augment Code has extended its Cosmos review system from code review to a full PR-to-merge loop, adding four capabilities: Verifier, PR Fixer, Review Dashboard, and cosmos approve. Dedicated Experts handle risk analysis, line-by-line correctness review, design review, runtime verification, and fixes.

    Why it matters: Augment has expanded code review into a PR-to-merge loop covering fixes, verification, and approval, giving readers a way to judge how multi-agent division of labor plays out in practice.

8/10Mon
  1. Hacker News · Claude Code 高分83

    Claude Code 将 auto mode 设为默认权限模式

    Anthropic 宣布从 8 月 14 日起,Pro、Max 和 Team 套餐的新会话默认运行 auto mode,并停止对分类器额外 token 开销收费;Enterprise、Claude API、AWS、Bedrock、Google Cloud 和 Microsoft Foundry 暂时保持可选,计划下个月改为默认。

    Awaiting translation

    Why it matters: Anthropic 公布 auto mode 的安全评测数据与内部拦截案例,可据此判断默认权限模式对现有工作流的影响。

8/5Wed
  1. Vercel · v0 Blog62

    Vercel ships v0 API for programmatic access to its app-generation agent

    Vercel ships v0 API, giving programmatic, headless access to the v0 app-generation agent: send a prompt, v0 generates an app, spins up a dev server in the Vercel Sandbox, and returns a preview URL you can embed in your own UI. The API is now generally available.

    Why it matters: v0 opens up its app-generation capability as an API, so readers can judge how to wire it into their own product or agent workflow.

7/30Thu
  1. Terminal-Bench · News62

    Terminal-Bench ships new Harbor features, turning the benchmark into a versioned asset that keeps getting updated

    The Terminal-Bench team has shipped a batch of new Harbor features that let datasets be released by version and let leaderboards migrate to new versions by reusing, re-evaluating, or rerunning trials. Tasks use semantic versioning: patch-level changes reuse old results as-is, validator changes only require re-evaluating saved artifacts, and only major changes that significantly alter the agent environment require a rerun. Dataset versions follow the highest version number among the tasks, and leaderboards use diffs to rerun only the tasks with major changes.

    Why it matters: The Terminal-Bench team maintains the benchmark like software, laying out concrete mechanisms for task semantic versioning and leaderboard upgrades that you can carry over to your own evaluation pipeline.

7/20Mon
  1. OpenAI Developer Blog · Codex71

    Codex Code Review now supports custom review rules in AGENTS.md

    OpenAI has added custom repository rules to Codex Code Review: you can put review guidelines in AGENTS.md, and Codex applies them during review and cites where each one came from in its findings. In OpenAI's own evaluation, the rule-guided version caught 98% of the required custom issues, versus 58.3% for the baseline. The guidance is to start with non-obvious invariants like compatibility requirements and data boundaries, put repo-level rules in the root directory and service-level rules in the corresponding directory, and leave formatting and mechanical checks to CI.

    Why it matters: OpenAI lays out the capabilities, the syntax, and the evaluation data for Codex Code Review custom rules, so you can judge how to bake your team's review experience into AGENTS.md.

7/15Wed