In February I hooked up 11 MCP Servers to my coding agent. Just the tool list in an empty session ate up 34,000 tokens, with Datadog alone contributing over two hundred tools. The agent got slower and kept picking the wrong tools.
Why it matters: I'm writing this up because 11 MCP Servers of my own blew up my context, and it showed me that tools and knowledge belong in different containers.
The author open-sourced claudemanager, a local daemon that Claude Code points to via ANTHROPIC_BASE_URL. It only changes the request's Authorization header to route sessions to the Max account with the most remaining capacity in its 5-hour, weekly, and per-model windows, switching at custom thresholds before those windows fill up.
Why it matters: The author also open-sourced a local proxy that automatically distributes Claude Code traffic across multiple Max accounts based on remaining quota, and logs requests along the way.
The author has open-sourced the Project Athena v9.9.9 kernel (MIT License), a local-first memory and governance framework built to solve the problems of Claude Code needing repeated corrections, sub-agents overstepping their bounds and modifying files, and losing state when it hits the quota limit mid-task.
Why it matters: The author validated a CLAUDE.md slimming and local-memory approach across 1900 sessions, and provides a directory structure and verification rules you can reuse directly.
The author has GPT (gpt-6.1-sol, reasoning tier high) handle planning, key decisions, and acceptance in Codex, and calls DeepSeek-V4.1-Flash through the official DeepSeek Harness to write code, run experiments, and fix bugs—together producing a local PDF toolkit with five working features.
Why it matters: The author splits the work between GPT for planning and DeepSeek for execution to get a PDF toolkit running, and shares three prompts plus cache usage data that can carry over to cutting costs on long tasks.
On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.
Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.
A product designer with six years of SaaS experience used Claude Code to single-handedly build Котомка, a life-planning app. Nearly all the code was written by AI; his job was to define requirements, review the results, and make decisions.
Why it matters: Using a real repository, the author documented the pitfalls he hit while building a product on Claude Code alone, plus the rules, hooks, and testing guardrails he set up around the AI.
Lovable launches Chats, an agent that runs at the workspace level, can hold conversations across projects and trigger builds; once changes are confirmed, it hands the task off to the project's builder agent and brings progress back into the conversation.
Why it matters: Lovable shares the three-layer architecture behind Chats—trajectory, inbox, and activation—which you can adapt for your own multi-agent orchestration.
Drawing on his own Codex setup, the author built a role-based model routing system for Claude Code: the main thread acts as coordinator, explorer uses Haiku for code search only, worker uses Opus for TDD implementation, verifier uses Sonnet to run checks independently, senior uses high-tier Opus for money, data, and concurrency, and reviewer switches to a different model for semantic review.
Why it matters: Based on a week of hands-on testing, the author shares the configuration, Hook enforcement, and cost trade-offs of multi-model division of labor in Claude Code, which you can adapt to your own multi-agent workflows.
In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.
Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.
Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.
Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.
Author Ryan Lopopolo argues that an agent is a parameterized program built on top of a set of capabilities: models and configurations, reasoning and tool-call loops, computers, disks, context, Skills, tools, connectors, runtimes, network policies, identity, IAM, guardrails, I/O channels, and system prompts.
Why it matters: Drawing on his experience building multiple agents, the author proposes a platform architecture that decouples capability interfaces from their implementations — a useful reference for teams building Agent platforms.
On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.
Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.
Drawing on a hands-on session where he dispatched 13 subagents to build a storyboard, the author walks through the subagents feature that both Claude Code and Codex have: subagents work in their own separate windows and hand only their conclusions back to the main conversation.
Why it matters: Using a hands-on session where he dispatched 13 subagents to build a storyboard, the author shows how subagents keep their work outside the main conversation and send back only the conclusions.
Anthropic has published a guide dissecting e-commerce AI agents. Drawing on deployment experience with retailers, e-commerce platforms, and teams in travel, entertainment, and telecom, it proposes a single-agent architecture that puts Claude in a standard agent loop, uses skills to cover long-tail needs, and calls tools to work with existing systems. The guide says that in comparative testing, this architecture beats both sub-agent designs and the approach of cramming everything into the prompt.
Why it matters: Drawing on enterprise e-commerce agent deployment experience, Anthropic lays out a complete engineering approach covering a single-agent-plus-skills architecture, latency and cost optimization, and memory and security evaluation.
Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.
Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.