Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.
Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.
OpenAI has open-sourced the harness that drives the Codex app, CLI, and IDE extensions, and through the Codex app-server client protocol it exposes capabilities like creating threads, starting turns, receiving events, and handling approval requests.
Why it matters: With the Codex harness and app-server protocol now public, developers can see how to embed the agent in their own products and where the boundaries are.
Addy Osmani describes how he runs 5 to 10 agents in parallel every day, usually with at most 5 running at once, and breaks loop engineering into two Claude Code primitives: goal, which drives a single bounded task toward a verifiable definition of done, and loop, which reruns a prompt at fixed intervals, much like cron.
Why it matters: The author breaks his daily routine of running 5 to 10 agents in parallel into two primitives—goal and loop—and shares reusable validation Skills and ways to write stop conditions.
8/11Tue
Tuesday
Paper Compute · Engineering BlogSelectedAI score8080
Spotify's platform team shared how the backend coding agent Honk evolved: it started out replacing migration scripts, then gradually took on build and test validation, raising merged PRs from 1000 per 3 months to 1000 per 10 days.
Why it matters: Spotify walks through Honk's evolution from script-based migration to a backend coding agent, focusing on how validation and standardization determine whether generated code can be merged.
The OpenAI Codex Cookbook lays out a set of repository conventions for wiring Codex into your development process: use AGENTS.md for persistent repository instructions, PLANS.md as the source of phase plans, and split the work into phased build files under harness/build/, with each phase spelling out its goals, acceptance criteria, boundaries, and approval gates.
Why it matters: OpenAI lays out a complete directory convention for constraining Codex with AGENTS.md, PLANS.md, and phased build files—one you can adapt to your own repository.
7/30Thu
Thursday
Martin Fowler · Exploring Generative AISelectedAI score7878
Thariq Shihipar, a member of Anthropic's engineering team, wrote up the new context engineering rules for Claude 5, saying the team has cut over 80% of the system prompt from Claude Code for models like Claude Opus 5 and Claude Fable 5, with no measurable loss on coding evals.
Augment Code proposes loop engineering: designing agent loops that run from trigger to execution to validation to outcome, with agents handling the intermediate steps and humans stepping in only at checkpoints that require judgment. The article compares loop engineering with prompt engineering and context engineering as distinct layers, lays out five stages—trigger, execution, validation, outcome, and improvement—and describes four team-level loops already running in production: code review, ticket-to-PR, vulnerability remediation, and incident response.
Why it matters: Augment Code breaks loop engineering into five stages—trigger, execution, validation, outcome, and improvement—and lays out four team-level loop patterns already running in production.
Addy Osmani proposes that a software factory has three layers—loop, harness, and factory. The factory isn’t a smarter agent; it’s multiple loops with harnesses feeding into a single review gate, with humans controlling the outer loop.
Why it matters: The author breaks the software factory into three layers—loop, harness, and factory—and points out that validation, not generation, is the real bottleneck.
Drawing on his experience during Stripe’s first code yellow, the author points out that after most code reds, all that’s left is a post-mortem and exhausted engineers, while the metrics go back to being unowned. He argues that a code red should leave behind a maintenance loop, and that OpenAI Codex’s /goal command can turn a one-off coding request into an ongoing objective with clear completion criteria, letting a persistent cluster of agents continuously watch metrics, generate interventions, and request human review.
Why it matters: Based on his Stripe code yellow experience, the author proposes using Codex’s /goal to turn one-off emergency fixes into a long-term maintenance loop that can carry over to SLO governance.
AI Hero's skills repo ships v1.1, renaming /to-prd to /to-spec, merging /to-plan and /to-issues into /to-tickets, and adding new Skills like /wayfinder, /research, and /prototype.
Why it matters: The author walks through the full Skill flow from grilling to deployment and gives the migration commands for the renames, the merge, and the new /wayfinder—useful for anyone building an AI development workflow.
Martin Fowler · Exploring Generative AISelectedAI score7171
The Claude Code team defines an AI agent loop as repeatedly running a work cycle until a predefined stopping condition is met, and splits it into four types—turn-based, goal-oriented, time-based, and proactive—along four dimensions: what triggers it, how it stops, the underlying instructions it uses, and the tasks it fits.
Why it matters: The Claude Code team breaks loops into four types—turn-based, goal-oriented, time-based, and proactive—and lays out how each is triggered, how it stops, and how tokens are controlled.
At Prime Radiant, author Jesse Vincent used Claude Code—working through the Slackline command-line Slack client—to collaborate with his own agent Ada: Claude proposes changes, Ada reviews and tests them, then Claude deploys the updates, forming a development loop where the agents review each other.
Why it matters: By looping two agents through mutual review, testing, and deployment, the author shows a transferable model for collaborative agent-based development.