Skip to content

#Context/memory

0 items today
3/13Fri
  1. Ryan Lopopolo60

    智能体利用率成为新的性能天花板

    作者 Ryan Lopopolo 提出智能体利用率是新的性能天花板,认为模型智能要转化为工程产出,取决于智能体在整个软件生命周期中被允许做多少事。他给出的数据是 2026 年 3 月使用 GPT-5.2 时,无 Symphony 每人每天 3-5 个 PR,接入 Symphony 后每人每周约 75 个 PR,差别来自 Symphony 让智能体保持忙碌。

    Awaiting translation

3/9Mon
  1. Paper Compute · Engineering Blog78

    我让 Agent 玩 1000 回合宝可梦,它始终没走出卧室

    作者构建的宝可梦红版自主 Agent 跑了 1000 回合仍停在卧室,原因是它把背景瓦片图地址 0xC4F2 当成文本框状态标志,一直按 A 键却不知道卡住。

    Awaiting translation

    Why it matters: 作者用 1000 回合卡在卧室的失败复盘,说明记录并回放 Agent 会话状态比改提示词更能定位静默故障。

3/8Sun
3/4Wed
2/23Mon
  1. OpenAI Developer Blog · Codex72

    用 Codex 跑 25 小时长时程任务:一份可复用的项目记忆文件栈

    OpenAI 用 GPT-5.3-Codex 在 Extra High 推理档下从空仓库连续运行约 25 小时、消耗约 13M token、生成约 3 万行代码,做出一个可测试的设计工具。

    Awaiting translation

    Why it matters: 作者用 25 小时、13M token 的实测展示长时程智能体如何靠持久化项目记忆和逐里程碑验证保持不跑偏。

2/17Tue
2/16Mon
  1. Hacker News · AGENTS.md62

    评估 AGENTS.md 对编程智能体是否有帮助

    一项研究评估了 AGENTS.md 这类上下文文件对编程智能体完成任务的帮助,发现提供上下文文件通常不会提升任务成功率,反而让推理成本平均增加超过 20%。该结论在不同大语言模型、编程智能体以及 LLM 生成和开发者提交的上下文文件上都成立。研究还发现智能体能较好遵循上下文文件中的指令,但被模型厂商推荐的仓库概览类内容并无帮助,作者建议任何试图提升性能的改动都应先经过严格评估再部署。

    Awaiting translation

2/11Wed
2/10Tue
  1. Paper Compute · Engineering Blog62

    Paper Compute 开源 tapes:为 AI 智能体提供透明遥测层

    Paper Compute 开源 tapes,一个位于 AI 智能体与推理服务之间的遥测层,用于为每次会话生成可持久保存、可审计的记录。它由代理服务、API server、CLI 客户端和终端 UI 四部分组成,代理服务无需改代码即可拦截并记录智能体与模型提供方之间的流量。

    Awaiting translation

  2. Jesse Vincent62

    用 Dorodango 比喻 AI 编程的日常打磨工作流

    作者 Jesse Vincent 把自己用 AI 构建软件的方式分成两类。一类是前期花大量时间做头脑风暴和规格文档,再让 Claude 或 Codex 生成实现计划并端到端验证,他称之为 fast waterfall;另一类是针对已有产品的小改动,打开产品、让 Claude 改、再看效果,属于日常打磨。

    Awaiting translation

2/9Mon
2/8Sun
  1. Martin Alderson78

    Automatically improving your CLAUDE.md file with agent session logs

    The author suggests using agent session logs to improve CLAUDE.md or AGENTS.md files: Claude Code stores sessions in ~/.claude/projects, while Codex stores them in ~/.codex/sessions—both in JSONL format but with different schemas.

    Why it matters: The author works backward from agent session logs to figure out what to improve in CLAUDE.md, and has open-sourced a CLI that cuts search time from several minutes down to seconds.

2/5Thu
  1. Martin Fowler · Exploring Generative AI75

    Context Engineering for Coding Agents: A Look at Configuration Options, Using Claude Code as an Example

    A Martin Fowler team article breaks down context engineering for coding agents, sorting context configuration into reusable prompts (instructions and guidelines), context interfaces (tools, MCP Servers, Skills), and workspace files. It then splits these by "who decides what gets loaded" into three categories: the LLM, the human, and the agent software.

    Why it matters: Using Claude Code as an example, this piece walks through how to configure context for coding agents and lays out the trade-offs between loading on demand and building up gradually.

1/30Fri
1/19Mon
  1. Kondasamy Jayaraman · Engineering Blog60

    RAG 入门:构建可用检索系统时没人告诉你的那些事

    作者结合自己搭建多个 RAG 系统的经验,把 RAG 拆成检索与生成两步,指出检索质量决定答案质量,提示词工程无法弥补糟糕的上下文。文中给出具体取舍:分块 300-500 token 并保留 20-30% 重叠,纯向量检索只能找到约 75% 的相关文档,混合检索可提升到 87%,重排能带来 20-35% 的准确率提升但增加 200-500ms 延迟。

    Awaiting translation

1/17Sat
  1. Geoffrey Huntley · Blog60

    Geoffrey Huntley:一切皆可成为 ralph 循环

    Geoffrey Huntley 认为软件开发方式已从像搭积木一样逐块纵向构建,转向以循环为单位来编排,他把自己定位为编程循环的工程师,而不是逐块写代码的人。他反对当下流行的多智能体与智能体间通信,理由是智能体本身是非确定性的,多智能体只会像微服务一样带来复杂度,而 ralph 是单仓库、单进程、每轮只做一件事的纵向扩展方案。

    Awaiting translation

1/8Thu
12/29Mon
12/27Sat
12/18Thu
12/17Wed
  1. Jesse Vincent66

    Claude Code's Skill not triggering? Maybe it never saw it at all.

    Claude Code lets the model know which Skills exist by injecting their names and descriptions into the system prompt. When there are too many Skills, or the description fields are too long, the system prompt stops listing them, so the model can't use them — and the prompt also tells the model not to use any Skill that isn't listed.

    Why it matters: The author explains why Claude Code doesn't trigger installed Skills, and gives a temporary fix using environment variables that you can apply right away.

12/2Tue
  1. Jesse Vincent69

    Building a front-end/back-end log bridge for the coding agent to make debugging web apps easier

    When developing web apps with the coding agent, the author often runs into client-side JavaScript bugs. If the agent can't fix them by reading the code, it fires up browser MCP for interactive debugging just to see the browser console logs—burning tokens and slowing things down.

    Why it matters: The author shares a reusable front-end/back-end log bridge approach that lets the coding agent see front-end logs without browser MCP.

11/16Sun
11/2Sun
10/27Mon
  1. OpenAI Developer Blog · Codex64

    Dagster Labs 如何用 Codex 做技术文档与教学

    Dagster Labs 分享了用 OpenAI Codex 加速技术文档写作、跨媒介内容转换和文档覆盖度评估的实践。他们重写了 CONTRIBUTING.md,明确文档层级、结构和最佳实践,让 Codex 能据此生成符合规范的文档;还借助 gh 命令让 Codex 解读 PR 的 diff 和描述,并让 Codex 把教程改写成 YouTube 视频脚本。

    Awaiting translation

    Why it matters: Dagster 团队把 Codex 用于文档写作、PR 解读和内容跨媒介转换,其中用文档生成代码来反向衡量文档覆盖度的做法可以迁移。

10/23Thu
  1. Jesse Vincent69

    Using episodic-memory to give Claude Code cross-session memory

    The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.

    Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.

10/19Sun
  1. Jesse Vincent71

    The author built a custom superpowers-chrome MCP that cut startup overhead from 13678 tokens to 947.

    The author built a lightweight Chrome MCP and Skill for Claude Code called superpowers-chrome. At startup, the MCP configuration takes up only 947 tokens, while Microsoft's Playwright MCP needs 13678 tokens just to be available—about 7% of the context window.

    Why it matters: The author compares the token overhead of a self-built Chrome MCP against Playwright MCP, laying out the concrete trade-offs involved in designing tool interfaces for LLMs.

10/15Wed
  1. Martin Fowler · Exploring Generative AI74

    拆解 Spec-Driven Development:Kiro、spec-kit 与 Tessl 三种工具实测

    Martin Fowler 试用 Kiro、spec-kit 和 Tessl 三款自称实现 spec-driven development(SDD)的工具,把 SDD 归纳为 spec-first、spec-anchored、spec-as-source 三个层次,并指出目前所有方案都停留在 spec-first。

    Awaiting translation

    Why it matters: 作者亲手试用 Kiro、spec-kit 和 Tessl 三款 SDD 工具,给出 spec-first、spec-anchored、spec-as-source 三层划分,并指出小任务被过度规格化的问题。

10/13Mon
  1. Jesse Vincent37

    让 AI 写出好文风的一个怪招:先让 Claude 读《The Elements of Style》

    开发者 Jesse Vincent 发现,让 Claude 先读 Strunk 1920 年版《The Elements of Style》再写 README,成稿比原来短约 30%,文风也更合他意。他把该书 HTML 转成约 12,000 词的 Markdown,因 Anthropic 的版权过滤机制拒绝处理这本已进入公有领域的书,最终改用 GPT-5 Codex 完成删减。

    Awaiting translation

10/1Wed
  1. Hacker News · Context Engineering 讨论76

    Anthropic on Effective Context Engineering for AI Agents

    Anthropic's applied AI team argues that context engineering is a continuation of prompt engineering, and the core idea is picking the smallest set of high-signal tokens within a limited attention budget.

    Why it matters: Anthropic lays out a systematic approach to context engineering, covering the trade-offs among three long-task strategies: compression, note-taking, and sub-agents.

9/29Mon
  1. Jesse Vincent66

    Rewriting CLAUDE.md Rules with GraphViz's Dot Language

    The author rewrote a long block of CLAUDE.md rules as a GraphViz dot flowchart, using quoted strings as node names, different shapes to distinguish decisions, commands, and warnings, and giving each flow an explicit trigger condition.

    Why it matters: After rewriting the CLAUDE.md rules as a GraphViz dot flowchart, Claude followed the rules better, and this approach can be carried over to your own projects.

9/25Thu
9/24Wed
  1. Hacker News · Context Engineering 讨论87

    Manus's AI Agent Context Engineering Experience: Six Practical Principles

    The Manus team shares context engineering lessons from building AI agents, centered on designing around the KV cache, managing tools by masking rather than removing them, treating the file system as context, steering attention by restating to-do items, keeping errors in context, and avoiding getting stuck on few-shot examples.

    Why it matters: The Manus team distilled lessons from rewriting their agent framework four times into six context engineering principles—ready to apply directly to your own agent implementation.

8/26Tue
7/9Wed
  1. Hacker News · Context Engineering 讨论62

    上下文工程实战:用 n8n 搭建多智能体深度研究工作流

    作者以自己用 n8n 搭建的多智能体深度研究应用为例,说明上下文工程不只是写提示词,而是设计并优化提供给 LLM 的完整上下文。他拆解了搜索规划智能体的系统提示词,涵盖指令、用户输入分隔符、子任务字段定义、结构化输出示例,以及注入当前日期时间让模型推断 start_date 和 end_date 的做法。

    Awaiting translation