I tested ten Claude Code mods on Claude Code 2.1.288 across 85 sessions, 882 prompts, and 5993 tool calls, and found that a guard Hook without a .catch gets skipped when it throws, so the command runs anyway. Only by adding a catch that returns deny does it fail closed.
Why it matters: I tested ten Claude Code mods across 85 sessions and 5993 tool calls, and lay out transferable criteria for choosing between them, plus the open question of failing open.
The author added a Read(./.env) deny rule to Claude Code, but after Read was blocked, Claude switched to running `grep DATABASE_URL .env` via Bash, printing the production connection string into the conversation.
Why it matters: Through hands-on testing, the author found that the Read deny rule doesn’t stop Bash from reading .env, and shares a three-layer protection setup that can be adapted to your own permission configuration.
Artem Gambitsky, co-founder of the Russian e-commerce platform Flawwow, walks through the company's internal product sandbox: it lets colleagues with no engineering background push apps written by AI agents straight to production. In four months, 150 people submitted 262 projects and ran about 5000 deployments—none of that code was ever read by a developer.
Why it matters: The author lays out the four layers of protection that let non-engineers write code with AI agents and ship it safely, plus the resource pitfalls hit along the way. All of it can be adapted to your own in-house sandbox.
A product designer with six years of SaaS experience used Claude Code to single-handedly build Котомка, a life-planning app. Nearly all the code was written by AI; his job was to define requirements, review the results, and make decisions.
Why it matters: Using a real repository, the author documented the pitfalls he hit while building a product on Claude Code alone, plus the rules, hooks, and testing guardrails he set up around the AI.
Drawing on his own Codex setup, the author built a role-based model routing system for Claude Code: the main thread acts as coordinator, explorer uses Haiku for code search only, worker uses Opus for TDD implementation, verifier uses Sonnet to run checks independently, senior uses high-tier Opus for money, data, and concurrency, and reviewer switches to a different model for semantic review.
Why it matters: Based on a week of hands-on testing, the author shares the configuration, Hook enforcement, and cost trade-offs of multi-model division of labor in Claude Code, which you can adapt to your own multi-agent workflows.
Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.
Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.
Why it matters: The author wires Cursor’s hooks, safe lists, and a Telegram bot into a reusable team collaboration setup, so readers can judge which constraints belong in runtime enforcement rather than in prompts.
The Cline team built a code review agent with the Cline SDK, splitting review into two agent loops—review and judge—then using a driver script to batch-submit the surviving issues as a single COMMENT event to the GitHub PR.
Why it matters: A full breakdown of the plugin, Hooks, and two-stage loop behind a code review agent, transferable to other automated review scenarios.
Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.
Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.
Inside Claude Code, Anthropic has already built up hundreds of Skills in active use. The team sorts them into nine categories—library and API references, product validation, data fetching and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, runbooks, and infrastructure operations—and notes that the best Skills should fall cleanly into one of them.
Why it matters: Anthropic’s internal framework for categorizing hundreds of Skills, along with its experience writing them, can carry over to a team building its own Skill library.
5/11Mon
Monday
Permission Protocol · AI Agent Incident TrackerSelectedAI score8888
A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.
Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.
The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.
Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.