Augment launches Project Builder on the Cosmos platform. This Cosmos expert turns a one-line feature description into a design doc grounded in the real codebase; after human review, it orchestrates worker agents to implement the work and drive it to merge.
Why it matters: Augment has shared how Project Builder handles design review and orchestration, plus the code volume and launch timelines of three production projects—enough to judge whether design-first plus agent orchestration is workable.
Inside Claude Code, Anthropic has already built up hundreds of Skills in active use. The team sorts them into nine categories—library and API references, product validation, data fetching and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, runbooks, and infrastructure operations—and notes that the best Skills should fall cleanly into one of them.
Why it matters: Anthropic’s internal framework for categorizing hundreds of Skills, along with its experience writing them, can carry over to a team building its own Skill library.
Anthropic has shipped dynamic workflows in Claude Code. Claude can write its own harness on the fly for a specific task, and these workflows can be shared and reused. Workflows orchestrate subagents through functions like agent(), parallel(), and pipeline(), and you can specify which model each agent uses and whether it runs in its own worktree. If a session is interrupted, resuming it picks up where it left off.
Why it matters: The Anthropic team breaks down six orchestration patterns for dynamic workflows and where each one fits, and these patterns carry over to multi-agent task design.
5/26Tue
Tuesday
Hacker News · AI Code Review 讨论SelectedAI score8888
Cloudflare built a CI-native AI code review system on top of the open-source coding agent OpenCode. A coordinating agent dispatches up to 7 dedicated review agents, split by security, performance, code quality, documentation, release, and internal standards, then deduplicates their output and posts a single structured review comment.
Why it matters: Cloudflare has published the plugin architecture, risk grading, and cost data behind its multi-agent code review in CI, and the setup can be ported to your own review pipeline.
Augment wires the Incident Investigator expert from its internal Cosmos platform into Slack and PagerDuty. It automatically triages every alert and runs root-cause analysis, then suggests one of four actions: fix the code, roll back, upgrade, or just keep monitoring. Humans only review the RCA and make the call.
Why it matters: Augment has shared the full playbook for putting Cosmos Expert on alert triage, along with a month of before-and-after data, so you can adapt it to your own on-call process.
Codex team member jason (@jxnlco) shares how to get the most out of Codex, the key being to combine persistent conversation threads, voice input, task intervention and queuing, MCP servers and connectors, conversation thread automation, goal setting, and the sidebar.
Why it matters: A member of the official Codex team breaks down how to use persistent conversation threads, task intervention, automation, and goal setting—approaches you can carry over into everyday agent workflows.
Starting with Codex 0.128.0, OpenAI offers Goals, turning one-off prompts into persistent objectives within a thread. Codex keeps checking evidence such as tests, benchmarks, or deliverables to decide whether the goal is done.
Why it matters: The official docs lay out where Goals fits, how to write its six elements, and the lifecycle commands, so you can tell when a persistent objective should replace a one-off prompt.
Claude Code team member Thariq makes the case for replacing Markdown with HTML as the output format for AI agents: HTML packs in more information, is easier to share, and supports two-way interaction, while Markdown's editing advantage stopped mattering once he switched to making changes through prompts.
Why it matters: Claude Code team members explain why they use HTML instead of Markdown as the agent output format, and share prompts you can use as-is along with the scenarios they fit.
Boris Cherny, the creator of Anthropic's Claude Code, said in an interview at Sequoia AI Ascent that he went all of 2026 without writing a single line of code, merging dozens of PRs a day—150 in a single day at his peak—doing most of his work from his phone, with 5 to 10 sessions and hundreds of agents running at any given time, plus thousands more chewing through deep tasks overnight.
Why it matters: Boris Cherny walks through Claude Code's path from incubation to a billion dollars in revenue, and lays out his calls: programming is solved, SaaS moats are being flattened, and organizational process is where the real edge is.