Augment Code pulled dozens of AGENTS.md files from its own monorepo and used its internal benchmark suite AuggieBench to compare how the same tasks performed with and without the file. The best files delivered a quality boost equivalent to upgrading from Haiku to Opus, while the worst made the output worse than having no AGENTS.md at all.
Why it matters: Augment Code used internal benchmarks to quantify how much the different ways of writing AGENTS.md actually differ, so readers can adjust their own repo's documentation structure accordingly.
After analyzing the architecture that surfaced in the Claude Code source leak, the author argues that its core isn't a secret algorithm but a while loop plus a tool dictionary in under 30 lines of Python, driven by stop_reason !
Why it matters: From the leaked source, the author distills 12 composable agent-engineering patterns and lays out a four-week path to get started, useful for checking your own implementation for gaps.
Baoyu walks through Claude Code's prompt caching mechanism to explain why quotas burn so fast, and lays out rules for saving tokens. He points out that caching only applies to prefixes, the main agent's cache window is 1 hour, and sub-agents' is 5 minutes. Reading from cache costs about one-tenth of recomputing, so frequent /clear actually triggers a full-price context rebuild. The rule of thumb: if the cache is still warm and the task hasn't changed, keep chatting; only start a new session when the cache has expired, the task has shifted, or there's too much context noise.
Why it matters: Starting from the prompt caching mechanism, this explains Claude Code's quota consumption and gives the criteria for deciding whether to continue a session or start over, plus configuration you can copy.
Based on the Claude Code source code that leaked unexpectedly last week, Drew Breunig mapped out how the system prompt is assembled: components fall into two categories—always included and conditionally included—and shift based on toggles like output_style, repl_mode, user_type_ant, skills_enabled, and mcp_connected.
Why it matters: The author breaks down the dynamic assembly logic behind Claude Code's system prompt, showing how conditional context engineering works in practice.
Thariq Shihipar, an engineer on Anthropic’s Claude Code team, summed up what the team learned from using hundreds of active Skills internally, sorting them into nine categories: library and API references, product validation, data acquisition and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, operations runbooks, and infrastructure operations.
Why it matters: Anthropic’s internal classification system for hundreds of Skills, along with its writing tips, can be adapted to help teams design their own Skills.
Paper Compute · Engineering BlogSelectedAI score7878
Bassim Eledath breaks the practical path of AI-assisted programming into 8 levels, from tab completion and agentic IDEs to context engineering, compound engineering, MCP and Skills, Harness Engineering, background agents, and finally autonomous agent teams.
Why it matters: The author lays out AI-assisted programming as 8 levels, from tab completion to autonomous agent teams, so readers can figure out where their own team stands.
The author used the open-source multimodal Qwen 3.5 series models for PDF OCR: first exporting each page as an image at 100 dpi with PyMuPDF, then feeding the images to the model for recognition. In testing, Qwen3.5-9B hit the sweet spot between quality and speed, while the smaller 0.8B to 2B models tended to go off track on complex documents, summarizing the content instead of transcribing it.
Why it matters: The author tested Qwen 3.5 models of various sizes for PDF OCR, and shares two reusable paths—local and via OpenRouter—along with cost data.
3/9Mon
Monday
Paper Compute · Engineering BlogSelectedAI score7878
A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.
Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.
1/27Tue
Tuesday
Martin Fowler · Exploring Generative AISelectedAI score7070