Skip to content

Context and memory

CLAUDE.md, AGENTS.md, context compression, and memory management: giving agents the information they need.

64 curated itemsRelated topicsWorkflowsCosts and usage limitsSkills

Latest curated items

Items 41–60 · 64 total
4/30Thu
  1. Augment Code · Blog71

    Augment Code put Karpathy-style rules to the test: the coding agent didn’t write better code, but it was cheaper and faster

    In AGENTS.md, Augment Code front-loads roughly 2.5k characters of Karpathy-style coding rules, then runs 40 OpenClaw PRs through Auggie, Claude Code, and Codex for comparison.

    Why it matters: A head-to-head test of three coding agents on the same set of PRs shows that prompt constraints mainly cut costs rather than improve quality, and it also surfaces differences between the harnesses.

4/22Wed
  1. Augment Code · Blog88

    Augment Code Tests AGENTS.md: A Good File Is Like a Model Upgrade, a Bad One Is Worse Than Nothing

    Augment Code pulled dozens of AGENTS.md files from its own monorepo and used its internal benchmark suite AuggieBench to compare how the same tasks performed with and without the file. The best files delivered a quality boost equivalent to upgrading from Haiku to Opus, while the worst made the output worse than having no AGENTS.md at all.

    Why it matters: Augment Code used internal benchmarks to quantify how much the different ways of writing AGENTS.md actually differ, so readers can adjust their own repo's documentation structure accordingly.

4/15Wed
  1. Kondasamy Jayaraman · Engineering Blog78

    12 Agent-Building Patterns Distilled from the Claude Code Source Leak

    After analyzing the architecture that surfaced in the Claude Code source leak, the author argues that its core isn't a secret algorithm but a while loop plus a tool dictionary in under 30 lines of Python, driven by stop_reason !

    Why it matters: From the leaked source, the author distills 12 composable agent-engineering patterns and lays out a four-week path to get started, useful for checking your own implementation for gaps.

4/10Fri
  1. Ryan Lopopolo65

    怎样才算把活干好:写清非功能性需求才能让 AI 智能体收敛

    Ryan Lopopolo 认为,AI 让验证问题变得明显,因为每个真实任务都依赖一个我们几乎从不写下来的问题,即怎样才算把活干好。产出和评审都涉及语气、品味、风险容忍度、打磨程度、可接受的捷径和完成标准等大量非功能性决策,过去团队靠组织设计、社交规范、招聘和入职把这些隐含规则传递给人,而模型无法走招聘流程,因此交给它的任务基本都欠规范。

    Awaiting translation

    Why it matters: 作者以在 OpenAI 做代码智能体的经历说明,非功能性需求不写下来,评审智能体就会陷入无休止的拉扯。

4/6Mon
  1. 宝玉78

    Claude Code Token-Saving Guide: Be Careful with the 1M Context—Neither Never Opening a New Session Nor Always Opening One Is Right

    Baoyu walks through Claude Code's prompt caching mechanism to explain why quotas burn so fast, and lays out rules for saving tokens. He points out that caching only applies to prefixes, the main agent's cache window is 1 hour, and sub-agents' is 5 minutes. Reading from cache costs about one-tenth of recomputing, so frequent /clear actually triggers a full-price context rebuild. The rule of thumb: if the cache is still warm and the task hasn't changed, keep chatting; only start a new session when the cache has expired, the task has shifted, or there's too much context noise.

    Why it matters: Starting from the prompt caching mechanism, this explains Claude Code's quota consumption and gives the criteria for deciding whether to continue a session or start over, plus configuration you can copy.

4/5Sun
  1. Drew Breunig78

    How Claude Code assembles system prompts

    Based on the Claude Code source code that leaked unexpectedly last week, Drew Breunig mapped out how the system prompt is assembled: components fall into two categories—always included and conditionally included—and shift based on toggles like output_style, repl_mode, user_type_ant, skills_enabled, and mcp_connected.

    Why it matters: The author breaks down the dynamic assembly logic behind Claude Code's system prompt, showing how conditional context engineering works in practice.

3/17Tue
  1. 宝玉87

    How Anthropic’s Team Uses Claude Code Skills: Nine Categories and Writing Tips

    Thariq Shihipar, an engineer on Anthropic’s Claude Code team, summed up what the team learned from using hundreds of active Skills internally, sorting them into nine categories: library and API references, product validation, data acquisition and analysis, business process automation, code scaffolding, code quality and review, CI/CD and deployment, operations runbooks, and infrastructure operations.

    Why it matters: Anthropic’s internal classification system for hundreds of Skills, along with its writing tips, can be adapted to help teams design their own Skills.

  2. Paper Compute · Engineering Blog78

    日志即自愈反馈回路:用遥测让智能体跨会话积累经验

    作者让智能体在 stereOS 虚拟机里用 PyBoy 无头运行宝可梦红,速度约为实时的 100 倍,智能体自己输出 NAV、BATTLE、BACKTRACK 等日志前缀,这些日志经 tapes 代理流入 Kafka,再由 Flink SQL 做 STUCK_LOOP、TOKEN_SPIKE 异常检测,JSONL 与 DuckDB 负责跨会话查询。

    Awaiting translation

    Why it matters: 作者用终端里跑宝可梦的智能体做实验,展示日志如何变成跨会话的观测记忆并反哺下一轮运行。

3/16Mon
  1. 宝玉78

    The 8 Levels of Agent Engineering: From Tab Completion to Autonomous Agent Teams

    Bassim Eledath breaks the practical path of AI-assisted programming into 8 levels, from tab completion and agentic IDEs to context engineering, compound engineering, MCP and Skills, Harness Engineering, background agents, and finally autonomous agent teams.

    Why it matters: The author lays out AI-assisted programming as 8 levels, from tab completion to autonomous agent teams, so readers can figure out where their own team stands.

3/9Mon
  1. Paper Compute · Engineering Blog78

    我让 Agent 玩 1000 回合宝可梦,它始终没走出卧室

    作者构建的宝可梦红版自主 Agent 跑了 1000 回合仍停在卧室,原因是它把背景瓦片图地址 0xC4F2 当成文本框状态标志,一直按 A 键却不知道卡住。

    Awaiting translation

    Why it matters: 作者用 1000 回合卡在卧室的失败复盘,说明记录并回放 Agent 会话状态比改提示词更能定位静默故障。

2/23Mon
  1. OpenAI Developer Blog · Codex72

    用 Codex 跑 25 小时长时程任务:一份可复用的项目记忆文件栈

    OpenAI 用 GPT-5.3-Codex 在 Extra High 推理档下从空仓库连续运行约 25 小时、消耗约 13M token、生成约 3 万行代码,做出一个可测试的设计工具。

    Awaiting translation

    Why it matters: 作者用 25 小时、13M token 的实测展示长时程智能体如何靠持久化项目记忆和逐里程碑验证保持不跑偏。

2/8Sun
  1. Martin Alderson78

    Automatically improving your CLAUDE.md file with agent session logs

    The author suggests using agent session logs to improve CLAUDE.md or AGENTS.md files: Claude Code stores sessions in ~/.claude/projects, while Codex stores them in ~/.codex/sessions—both in JSONL format but with different schemas.

    Why it matters: The author works backward from agent session logs to figure out what to improve in CLAUDE.md, and has open-sourced a CLI that cuts search time from several minutes down to seconds.

2/5Thu
  1. Martin Fowler · Exploring Generative AI75

    Context Engineering for Coding Agents: A Look at Configuration Options, Using Claude Code as an Example

    A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.

    Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.

1/30Fri
12/17Wed
  1. Jesse Vincent66

    Claude Code's Skill not triggering? Maybe it never saw it at all.

    Claude Code lets the model know which Skills exist by injecting their names and descriptions into the system prompt. When there are too many Skills, or the description fields are too long, the system prompt stops listing them, so the model can't use them — and the prompt also tells the model not to use any Skill that isn't listed.

    Why it matters: The author explains why Claude Code doesn't trigger installed Skills, and gives a temporary fix using environment variables that you can apply right away.

12/2Tue
  1. Jesse Vincent69

    Building a front-end/back-end log bridge for the coding agent to make debugging web apps easier

    When developing web apps with the coding agent, the author often runs into client-side JavaScript bugs. If the agent can't fix them by reading the code, it fires up browser MCP for interactive debugging just to see the browser console logs—burning tokens and slowing things down.

    Why it matters: The author shares a reusable front-end/back-end log bridge approach that lets the coding agent see front-end logs without browser MCP.

10/27Mon
  1. OpenAI Developer Blog · Codex64

    Dagster Labs 如何用 Codex 做技术文档与教学

    Dagster Labs 分享了用 OpenAI Codex 加速技术文档写作、跨媒介内容转换和文档覆盖度评估的实践。他们重写了 CONTRIBUTING.md,明确文档层级、结构和最佳实践,让 Codex 能据此生成符合规范的文档;还借助 gh 命令让 Codex 解读 PR 的 diff 和描述,并让 Codex 把教程改写成 YouTube 视频脚本。

    Awaiting translation

    Why it matters: Dagster 团队把 Codex 用于文档写作、PR 解读和内容跨媒介转换,其中用文档生成代码来反向衡量文档覆盖度的做法可以迁移。

10/23Thu
  1. Jesse Vincent69

    Using episodic-memory to give Claude Code cross-session memory

    The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.

    Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.

10/19Sun
  1. Jesse Vincent71

    The author built a custom superpowers-chrome MCP that cut startup overhead from 13678 tokens to 947.

    The author built a lightweight Chrome MCP and Skill for Claude Code called superpowers-chrome. At startup, the MCP configuration takes up only 947 tokens, while Microsoft's Playwright MCP needs 13678 tokens just to be available—about 7% of the context window.

    Why it matters: The author compares the token overhead of a self-built Chrome MCP against Playwright MCP, laying out the concrete trade-offs involved in designing tool interfaces for LLMs.

10/15Wed
  1. Martin Fowler · Exploring Generative AI74

    拆解 Spec-Driven Development:Kiro、spec-kit 与 Tessl 三种工具实测

    Martin Fowler 试用 Kiro、spec-kit 和 Tessl 三款自称实现 spec-driven development(SDD)的工具,把 SDD 归纳为 spec-first、spec-anchored、spec-as-source 三个层次,并指出目前所有方案都停留在 spec-first。

    Awaiting translation

    Why it matters: 作者亲手试用 Kiro、spec-kit 和 Tessl 三款 SDD 工具,给出 spec-first、spec-anchored、spec-as-source 三层划分,并指出小任务被过度规格化的问题。