Thariq Shihipar, a member of Anthropic's engineering team, wrote up the new context engineering rules for Claude 5, saying the team has cut over 80% of the system prompt from Claude Code for models like Claude Opus 5 and Claude Fable 5, with no measurable loss on coding evals.
A QA engineer distilled six months of experience doing web testing on Claude Code into an open-source Skill package called paranoid-qa. At its core is an evidence contract: Pass/Fail can only be based on actual artifacts like screenshots, request bodies, and logs; anything unverified gets marked Not tested; anything blocked by the environment gets marked Blocked; and forms must verify the real submitted payload.
Why it matters: The author codified six months of QA experience into a testing Skill package for Claude Code, along with an evidence contract and failure checklist that can be reused directly.
The Claude Code team defines an AI agent loop as repeatedly running a work cycle until a predefined stopping condition is met, and splits it into four types—turn-based, goal-oriented, time-based, and proactive—along four dimensions: what triggers it, how it stops, the underlying instructions it uses, and the tasks it fits.
Why it matters: The Claude Code team breaks loops into four types—turn-based, goal-oriented, time-based, and proactive—and lays out how each is triggered, how it stops, and how tokens are controlled.
At Prime Radiant, author Jesse Vincent used Claude Code—working through the Slackline command-line Slack client—to collaborate with his own agent Ada: Claude proposes changes, Ada reviews and tests them, then Claude deploys the updates, forming a development loop where the agents review each other.
Why it matters: By looping two agents through mutual review, testing, and deployment, the author shows a transferable model for collaborative agent-based development.
6/8Mon
Monday
Permission Protocol · AI Agent Incident TrackerSelectedAI score8888
Inside Claude Code, Anthropic has already built up hundreds of Skills in active use. The team sorts them into nine categories—library and API references, product validation, data fetching and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, runbooks, and infrastructure operations—and notes that the best Skills should fall cleanly into one of them.
Why it matters: Anthropic’s internal framework for categorizing hundreds of Skills, along with its experience writing them, can carry over to a team building its own Skill library.
Permission Protocol · AI Agent Incident TrackerSelectedAI score8080
Anthropic has shipped dynamic workflows in Claude Code. Claude can write its own harness on the fly for a specific task, and these workflows can be shared and reused. Workflows orchestrate subagents through functions like agent(), parallel(), and pipeline(), and you can specify which model each agent uses and whether it runs in its own worktree. If a session is interrupted, resuming it picks up where it left off.
Why it matters: The Anthropic team breaks down six orchestration patterns for dynamic workflows and where each one fits, and these patterns carry over to multi-agent task design.
5/27Wed
Wednesday
Permission Protocol · AI Agent Incident TrackerSelectedAI score7878
The Claude Code CLI has a critical RCE vulnerability: an attacker can craft a claude-cli:// deeplink to exploit eagerParseCliFlag's context-free parsing of process.argv in main.tsx.
Why it matters: I walked through the RCE chain caused by Claude Code's lack of contextual parsing for command-line arguments, and gave my take on where the authorization boundary should be drawn.
Permission Protocol · AI Agent Incident TrackerSelectedAI score7878