用 scoped .mdc globs 解决 Cursor Composer 的 token 膨胀和响应变慢
Cursor Composer 变慢、token 消耗快,常见原因是多个 .mdc 文件设了 alwaysApply: true,4 到 5 条规则会给每轮对话前置 3000+ token 的静态指令,到第 10 轮时 prompt prefill 延迟明显上升。
Awaiting translation
Cursor Composer 变慢、token 消耗快,常见原因是多个 .mdc 文件设了 alwaysApply: true,4 到 5 条规则会给每轮对话前置 3000+ token 的静态指令,到第 10 轮时 prompt prefill 延迟明显上升。
Awaiting translation
一项实证研究用固定执行循环的轻量编码 harness,在 SWE-Bench Verified 和 Terminal-Bench 2.1 上对四个模型评测 176 组匹配设置,变量为规划、动作空间和上下文管理。
Awaiting translation
作者指出把架构规范和 lint 要求全塞进单个 .cursorrules 会让 Cursor Composer 每轮注入约 2000 token,10 轮对话就重复消耗 20000 token,并挤占上下文窗口导致 Composer 提前丢弃旧消息和文件 diff。
Awaiting translation
文章对比 Claude Code v2.1.274、Codex CLI rust-v0.154.0 和 Gemini CLI v0.60.0 在 monorepo 中查找指令文件的规则:三者都沿目录链拼接而非覆盖,但起点、每目录文件数和总量上限不同。
Awaiting translation
作者认为 Cursor Composer 在新会话或长对话后违反架构规则,原因是上下文窗口淘汰和扁平化提示词衰减,而非模型变笨。他给出的做法是把 .cursorrules 写成带身份指令、技术栈、负面约束、输出前检查清单和恢复指令的分层结构,并在多目录仓库中按 backend/、frontend/ 拆分作用域规则,让就近目录的规则获得更高权重。
Awaiting translation
作者认为前沿模型出错多是因为对目标或优先级做了错误假设,而不是理解不了任务,因此提示词应交代背景和优先级,而非只给具体规格。他给出自己给 Deckard 写的提示词作为示例,其中约一半内容在说明长期目标、个人使用场景和保持笔记本不烫等优先级,结果模型提出了用 native messaging 运行本地模型、改用 Gradient 模型等规格里没写的改进。
Awaiting translation
一位创业公司首位工程招聘在 Hacker News 发问:如何把遗留代码库改造得更易被 AI 编程智能体理解。他所在的 2 年历史初创公司代码库始于 CTO 用 Lovable 搭的 MVP,如今存在业务逻辑重复、僵尸表和列、隐式依赖等问题。他计划把整个代码库文档化为 AI 智能体的“知识数据库”,并参考 Meta 的一篇文章,希望听取有类似经历的工程师分享做法、结果与避坑建议。
Awaiting translation
作者分享自己在项目中使用的提示词:让智能体在当前仓库落地一套最小化的 Google Open Knowledge Format(OKF v0.2)工程知识库,并把使用规则写进根目录 AGENTS.md。
Awaiting translation
作者对比 Google 的 Open Knowledge Format(OKF)与 OpenViking 两种智能体记忆方案,认为二者解决相似问题但层次不同。
Awaiting translation
文章按厂商文档把 Claude Code v2.1.270、Codex CLI rust-v0.154.0 和 Gemini CLI v0.59.0 的跨会话记忆拆成三层:可恢复的会话记录、上下文填满后替换历史的摘要、以及智能体自己写入的记忆。
Awaiting translation
作者把记忆 MCP server 接入 Cursor 后看到绿色状态,但智能体并不会自动记住东西,因为连接只提供了工具,还需要规定何时调用。
Awaiting translation
Sean Goedecke 认为,大多数“为 AI 智能体打造 X”的尝试会失败,理由有三:适合智能体的工具同样适合人类,智能体像工程师一样输入文本、调用 API、读图和分类;现有工具已存在于训练数据中,新工具若只比人类版好 20% 也难被采用,这也是他不看好为智能体设计新编程语言的原因;智能体的理想工效学尚无定论,静态类型语言等说法都能讲出正反两套故事。
Awaiting translation
有开发者反馈 Claude Code 项目开发数月后开始频繁触发压缩会话(Compacting conversation),压缩完记忆就错乱,询问如何迁移必要记忆。
Awaiting translation
In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.
Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.
Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.
Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.
Stolen Thoughts 研究发现 OpenAI、Anthropic 和 Google 的 reasoning API 存在漏洞:加密 reasoning block 未与具体模型、会话和用户充分绑定,把强模型的加密 reasoning 传给同厂商弱模型并越狱后,弱模型会以明文输出强模型的推理内容。
Awaiting translation
Why it matters: 研究揭示加密 reasoning block 可跨模型解密,并给出公开轨迹中泄露密钥的实测数据,对智能体基础设施设计有直接参考价值。
作者长期高强度使用 OpenCode + OpenChamber + OpenCode Go 与 Ollama Cloud 的 DeepSeek V4 Flash 订阅,同时跑 15 个项目额度消耗很少。
Awaiting translation
作者给出一份可直接复用的 AGENTS.md 仓库规则模板,用于给编程智能体提供仓库级指令,涵盖根目录文件、子目录嵌套覆盖(AGENTS.md 或 AGENTS.override.md)、精确的验证命令、文件归属、部署规则和需要停止的操作。
Awaiting translation
OKF Agent Memory 是一个基于 Open Knowledge Format(OKF)v0.2 的 Git 原生持久项目记忆工具,用双层记忆架构解决上下文窗口关闭后对话重置、架构决策丢失的问题。
Awaiting translation
Author Ryan Lopopolo argues that an agent is a parameterized program built on top of a set of capabilities: models and configurations, reasoning and tool-call loops, computers, disks, context, Skills, tools, connectors, runtimes, network policies, identity, IAM, guardrails, I/O channels, and system prompts.
Why it matters: Drawing on his experience building multiple agents, the author proposes a platform architecture that decouples capability interfaces from their implementations — a useful reference for teams building Agent platforms.
文章对比了 CLAUDE.md、AGENTS.md、.cursor/rules 和 GEMINI.md 四种上下文文件约定,说明各自的加载顺序、合并方式与体积上限。
Awaiting translation
文章在 commit 5dc9490b 上阅读 Aider 源码,解释 tree-sitter repo map 如何按依赖图排序并在 --map-tokens(默认 1k token)预算内裁剪,以及 whole、diff、diff-fenced、udiff、editor-diff/editor-whole 各编辑格式分别适配哪类模型。
Awaiting translation
作者提出用六段式 spec 模板(Goal、Non-goals、Interfaces、Files、Verification、Budget)向编码智能体描述任务,并配了一个在交给智能体前检查 spec 的 linter。
Awaiting translation
Why it matters: 给出可直接套用的六段式 spec 模板、tally 实例和配套 linter,读者能据此改造自己交给编码智能体的任务描述。
作者按 commit f7fb0c4b 阅读 OpenHands 仓库,梳理其架构:Agent Canvas 是托管编码智能体的控制中心,真正的智能体在 Software Agent SDK 中,采用单步执行模型,每步依次检查待确认动作、压缩事件历史、查询 LLM、解析响应、确认门禁、执行工具并生成 ObservationEvent。
Awaiting translation
文章把 coding-agent harness 定义为模型外层的循环,并拆成四部分:循环、工具、权限门禁和上下文,每部分都对照真实实现核对。
Awaiting translation
作者用三款 Claude 模型在多个 Python 和 TypeScript 仓库上对比 grep 词法检索与 LSP 语义导航,发现简单代码定位任务中模型只有 0% 到 6% 会选择语义工具,强制走语义路径反而让成功率从 100% 降到 89%。
Awaiting translation
Claude Code 的实验性功能 Agent Teams 让多个独立 AI 实例在本地终端里互发消息、开会辩论,与子代理的上下级关系不同,团队成员是平级点对点通信。
Awaiting translation
On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.
Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.
GitHub Copilot 团队复盘了四项降低 AI 编码成本的改动,核心原则是按完整任务而非单次工具调用衡量效率。
Awaiting translation
Why it matters: GitHub Copilot 团队复盘四项降本改动,并给出可迁移的评估方法:按完整任务而非单次工具调用衡量成本。
Drawing on a hands-on session where he dispatched 13 subagents to build a storyboard, the author walks through the subagents feature that both Claude Code and Codex have: subagents work in their own separate windows and hand only their conclusions back to the main conversation.
Why it matters: Using a hands-on session where he dispatched 13 subagents to build a storyboard, the author shows how subagents keep their work outside the main conversation and send back only the conclusions.
Anthropic has published a guide dissecting e-commerce AI agents. Drawing on deployment experience with retailers, e-commerce platforms, and teams in travel, entertainment, and telecom, it proposes a single-agent architecture that puts Claude in a standard agent loop, uses skills to cover long-tail needs, and calls tools to work with existing systems. The guide says that in comparative testing, this architecture beats both sub-agent designs and the approach of cramming everything into the prompt.
Why it matters: Drawing on enterprise e-commerce agent deployment experience, Anthropic lays out a complete engineering approach covering a single-agent-plus-skills architecture, latency and cost optimization, and memory and security evaluation.
一位 Claude Code 工程师写的《找到你的未知》长文四天获得三百多万浏览,核心是把提示词看作地图、代码库看作地盘,两者之间的差距就是未知,模型越强,卡住它的越可能是使用者自己讲不清未知。
Awaiting translation
作者用自己在用的 Remotion 视频工程演示 CLAUDE.md 的分层写法:根目录 CLAUDE.md 只放路由,41 行,指向 L0 官方、共享、出片三层资源;产出目录的 CLAUDE.md 同样只做路由,61 行。
Awaiting translation
Anthropic 工程师在《Claude 5 时代模型的上下文工程新规则》中称,针对 Opus 5、Fable 5 这一代模型把 Claude Code 系统提示词删掉 80% 以上,跑官方评测没有可测出的损失。
Awaiting translation
Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.
Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.
Semantic Overlays 在冻结模型上训练小型适配器,把一段 token 在残差流中标记为“不可执行”,文本仍可读但失去下达指令的权限,且没有文本能模仿这个标记。
Awaiting translation
模型能力再强,其"好工作"的先验也未必与你的标准一致,因此始终需要围绕模型整理环境,让它在非功能性要求上偏向你或组织认可的选择。模型变好不会让这种上下文学习(ICL)需求消失。所谓 Harness Engineering,本质是一组技巧,为 ICL 提供即时机会,在不束缚推理模型的前提下对齐模型行为。
Awaiting translation
宝玉在演讲中提出 AI 原生思维,主张做 AI 产品要盯着模型能力边界线找需求,并按能力、成本、价值三条边界判断值不值得做。他以自己做的字幕翻译 App BaoCut 为例,说明从模拟字幕组的 V1 转向以终为始的 V2 后,用词级时间戳对齐、术语表注入和 Agent 自验证替代人工校对,一次成本优化把调用从 33 次降到 12 次、单集处理从 31 分钟降到 18 分钟。
Awaiting translation
Addy Osmani 在博客中提出,智能体能跳过写代码、调试、读别人代码这些原本积累经验的环节,因此刻意练习变得必要,他建议在提示前先形成假设、多问为什么、读 diff、预测失败点,并偶尔手写小问题。
Awaiting translation
作者提出用一段提示词让编码智能体就当前代码库出 5 道难度递增的题,答完后由它解释自己理解错的地方,提示词中要求使用 askuserquestiontool。作者认为这种方式比让智能体直接描述代码更有效,因为它掌握了自己所想与代码实际行为的差距;在复杂重构前先让智能体就方案出题,也能提前发现问题,同样的做法还可用于复杂表格或文档集,甚至可以要求通过 PR 测验后才允许智能体创建 PR。
Awaiting translation