In AGENTS.md, Augment Code front-loads roughly 2.5k characters of Karpathy-style coding rules, then runs 40 OpenClaw PRs through Auggie, Claude Code, and Codex for comparison.
Why it matters: A head-to-head test of three coding agents on the same set of PRs shows that prompt constraints mainly cut costs rather than improve quality, and it also surfaces differences between the harnesses.
Augment Code pulled dozens of AGENTS.md files from its own monorepo and used its internal benchmark suite AuggieBench to compare how the same tasks performed with and without the file. The best files delivered a quality boost equivalent to upgrading from Haiku to Opus, while the worst made the output worse than having no AGENTS.md at all.
Why it matters: Augment Code used internal benchmarks to quantify how much the different ways of writing AGENTS.md actually differ, so readers can adjust their own repo's documentation structure accordingly.
After analyzing the architecture that surfaced in the Claude Code source leak, the author argues that its core isn't a secret algorithm but a while loop plus a tool dictionary in under 30 lines of Python, driven by stop_reason !
Why it matters: From the leaked source, the author distills 12 composable agent-engineering patterns and lays out a four-week path to get started, useful for checking your own implementation for gaps.
Ryan Lopopolo 认为,AI 让验证问题变得明显,因为每个真实任务都依赖一个我们几乎从不写下来的问题,即怎样才算把活干好。产出和评审都涉及语气、品味、风险容忍度、打磨程度、可接受的捷径和完成标准等大量非功能性决策,过去团队靠组织设计、社交规范、招聘和入职把这些隐含规则传递给人,而模型无法走招聘流程,因此交给它的任务基本都欠规范。
Baoyu walks through Claude Code's prompt caching mechanism to explain why quotas burn so fast, and lays out rules for saving tokens. He points out that caching only applies to prefixes, the main agent's cache window is 1 hour, and sub-agents' is 5 minutes. Reading from cache costs about one-tenth of recomputing, so frequent /clear actually triggers a full-price context rebuild. The rule of thumb: if the cache is still warm and the task hasn't changed, keep chatting; only start a new session when the cache has expired, the task has shifted, or there's too much context noise.
Why it matters: Starting from the prompt caching mechanism, this explains Claude Code's quota consumption and gives the criteria for deciding whether to continue a session or start over, plus configuration you can copy.
Based on the Claude Code source code that leaked unexpectedly last week, Drew Breunig mapped out how the system prompt is assembled: components fall into two categories—always included and conditionally included—and shift based on toggles like output_style, repl_mode, user_type_ant, skills_enabled, and mcp_connected.
Why it matters: The author breaks down the dynamic assembly logic behind Claude Code's system prompt, showing how conditional context engineering works in practice.
Thariq Shihipar, an engineer on Anthropic’s Claude Code team, summed up what the team learned from using hundreds of active Skills internally, sorting them into nine categories: library and API references, product validation, data acquisition and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, operations runbooks, and infrastructure operations.
Why it matters: Anthropic’s internal classification system for hundreds of Skills, along with its writing tips, can be adapted to help teams design their own Skills.
Paper Compute · Engineering BlogSelectedAI score7878
Bassim Eledath breaks the practical path of AI-assisted programming into 8 levels, from tab completion and agentic IDEs to context engineering, compound engineering, MCP and Skills, Harness Engineering, background agents, and finally autonomous agent teams.
Why it matters: The author lays out AI-assisted programming as 8 levels, from tab completion to autonomous agent teams, so readers can figure out where their own team stands.
3/9Mon
Monday
Paper Compute · Engineering BlogSelectedAI score7878
The author suggests using agent session logs to improve CLAUDE.md or AGENTS.md files: Claude Code stores sessions in ~/.claude/projects, while Codex stores them in ~/.codex/sessions—both in JSONL format but with different schemas.
Why it matters: The author works backward from agent session logs to figure out what to improve in CLAUDE.md, and has open-sourced a CLI that cuts search time from several minutes down to seconds.
2/5Thu
Thursday
Martin Fowler · Exploring Generative AISelectedAI score7575
A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.
Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.
Claude Code lets the model know which Skills exist by injecting their names and descriptions into the system prompt. When there are too many Skills, or the description fields are too long, the system prompt stops listing them, so the model can't use them — and the prompt also tells the model not to use any Skill that isn't listed.
Why it matters: The author explains why Claude Code doesn't trigger installed Skills, and gives a temporary fix using environment variables that you can apply right away.
When developing web apps with the coding agent, the author often runs into client-side JavaScript bugs. If the agent can't fix them by reading the code, it fires up browser MCP for interactive debugging just to see the browser console logs—burning tokens and slowing things down.
The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.
Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.
The author built a lightweight Chrome MCP and Skill for Claude Code called superpowers-chrome. At startup, the MCP configuration takes up only 947 tokens, while Microsoft's Playwright MCP needs 13678 tokens just to be available—about 7% of the context window.
Why it matters: The author compares the token overhead of a self-built Chrome MCP against Playwright MCP, laying out the concrete trade-offs involved in designing tool interfaces for LLMs.
10/15Wed
Wednesday
Martin Fowler · Exploring Generative AISelectedAI score7474