The author instrumented a demo ticket-booking MCP server with a self-built tool called mcpspan and identified five kinds of agent invocation problems that never show up in logs: agents guessing tool names that don't exist; parameters that don't match the schema and get rejected by the SDK before the handler runs; parameter types misunderstood because of how the tool description is worded; retry loops that keep hitting the same parameter; and responses so large they eat into the context window.
Why it matters: The author used a self-built MCP analysis tool to empirically surface five failure modes in agent calls that logs don't reveal, and these can be adapted to troubleshoot your own MCP server.
The author argues that prompt injection shouldn't be solved by making models smarter; instead, just as SQL injection is handled with parameterized queries, the boundary should be drawn where the agent executes actions.
Why it matters: The author draws an analogy between prompt injection and SQL injection, arguing for moving the boundary from the model to the tool-call layer, and lays out a practical approach with a strategy layer and separate read and write phases.
In February I hooked up 11 MCP Servers to my coding agent. Just the tool list in an empty session ate up 34,000 tokens, with Datadog alone contributing over two hundred tools. The agent got slower and kept picking the wrong tools.
Why it matters: I'm writing this up because 11 MCP Servers of my own blew up my context, and it showed me that tools and knowledge belong in different containers.
Pennyforge ran anonymous health checks on the 78 servers that responded to initialize out of 186 endpoints in the a–b slice of a public MCP registry. Only 40 of them (51.3%) made it through the full flow of initialize → tools/list → one safe tools/call.
Flash-Agents is a Claude Code plugin that delegates bounded coding work—implementing slices, porting tests, reviewing diffs, mapping out a codebase—to a DeepSeek V4.1 Flash worker, while Claude keeps architecture, acceptance criteria, and final review.
Why it matters: The author outsources Claude Code's coding tasks to a DeepSeek Flash worker and shares the sandbox, patches, and measured data, so you can judge the cost and safety boundaries for yourself.
The author has open-sourced the Project Athena v9.9.9 kernel (MIT License), a local-first memory and governance framework built to solve the problems of Claude Code needing repeated corrections, sub-agents overstepping their bounds and modifying files, and losing state when it hits the quota limit mid-task.
Why it matters: The author validated a CLAUDE.md slimming and local-memory approach across 1900 sessions, and provides a directory structure and verification rules you can reuse directly.
RugSnare is a runtime integrity tool for MCP tool descriptions. It computes a normalized hash pin over each approved tool's { name, description, inputSchema }, and any silent change afterward triggers an alert and fails CI (exit 1).
Why it matters: RugSnare hash-pins MCP tool descriptions, keeps watching for silent changes after approval, and shares measured data from 66 official server versions.
On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.
Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.
9/29Tue
Tuesday
Simon Willison · Coding AgentsSelectedAI score7878
Simon Willison live-blogged the OpenAI DevDay 2026 keynote from Fort Mason in San Francisco, where OpenAI announced the personal agent Dots, ChatGPT Space, GPT-6.1 Sol, Ultrafast, and more.
Why it matters: A running, item-by-item record of what OpenAI announced at DevDay, for a quick look at what Dots, GPT-6.1 Sol, Ultrafast, and Codex Security actually look like.
The open-source project Paseo pulls command-line agents like Claude Code, Codex, OpenCode, and Pi into a single management interface. All agents still run locally on your own machine, and the project has already reached 18700+ stars.
Why it matters: The original article shows how to bring multiple command-line agents into one interface and hand off tasks across agents, making it a useful reference for developers running several agents at once.
Why it matters: The author tested 10 open-source detectors against 629 real injection attacks, with full comparison data at both default thresholds and after calibration.
Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.
Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.
Why it matters: The author wires Cursor’s hooks, safe lists, and a Telegram bot into a reusable team collaboration setup, so readers can judge which constraints belong in runtime enforcement rather than in prompts.
Author Ryan Lopopolo argues that an agent is a parameterized program built on top of a set of capabilities: models and configurations, reasoning and tool-call loops, computers, disks, context, Skills, tools, connectors, runtimes, network policies, identity, IAM, guardrails, I/O channels, and system prompts.
Why it matters: Drawing on his experience building multiple agents, the author proposes a platform architecture that decouples capability interfaces from their implementations — a useful reference for teams building Agent platforms.
Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.
Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.
The author used a targeted prompt injection attack chain to reach a 60-80% attack success rate in Claude Code Opus 5 Auto Mode (small sample), whereas a third-party evaluation commissioned by Anthropic had reported a 0.00% injection success rate.
Why it matters: The author used a module-obscuring attack chain to reach a 60-80% success rate in Auto Mode, showing that the classifier is not a sandbox.