Augment launches Project Builder on the Cosmos platform. This Cosmos expert turns a one-line feature description into a design doc grounded in the real codebase; after human review, it orchestrates worker agents to implement the work and drive it to merge.
Why it matters: Augment has shared how Project Builder handles design review and orchestration, plus the code volume and launch timelines of three production projects—enough to judge whether design-first plus agent orchestration is workable.
Terminal-Bench has released version 2.0 and the Harbor package. The former is a more rigorously validated, harder benchmark for evaluating agents; the latter is for evaluating and optimizing agents. Harbor rewrites Terminal-Bench's test harness, supports deploying containers in the cloud, provides rollout interfaces for RL and SFT, and works with any agent you can put in a container.
Why it matters: Terminal-Bench 2.0 and Harbor are released together, so readers can see how the agent evaluation benchmark is validated and how to scale it in the cloud.
Anthropic rolled out its first-party Skills system simultaneously on Claude Code, Claude.ai, and the Claude API, and author Jesse Vincent quickly followed with a new version of Superpowers built on the official Skills.
Why it matters: Drawing on nearly a month of hands-on use, the author compares the official Skills with his own setup and lays out the trade-offs involved in migrating.