Skip to content

#Context/memory

0 items today
9/19Sat
9/18Fri
9/17Thu
  1. Cursor Forum · Guides71

    用分层 .cursorrules 防止 Cursor Composer 跨会话重启后忽略架构规则

    作者认为 Cursor Composer 在新会话或长对话后违反架构规则,原因是上下文窗口淘汰和扁平化提示词衰减,而非模型变笨。他给出的做法是把 .cursorrules 写成带身份指令、技术栈、负面约束、输出前检查清单和恢复指令的分层结构,并在多目录仓库中按 backend/、frontend/ 拆分作用域规则,让就近目录的规则获得更高权重。

    Awaiting translation

9/15Tue
  1. Sean Goedecke · Blog66

    给 AI 智能体讲清目标优先级,而不只是具体做法

    作者认为前沿模型出错多是因为对目标或优先级做了错误假设,而不是理解不了任务,因此提示词应交代背景和优先级,而非只给具体规格。他给出自己给 Deckard 写的提示词作为示例,其中约一半内容在说明长期目标、个人使用场景和保持笔记本不烫等优先级,结果模型提出了用 native messaging 运行本地模型、改用 Gradient 模型等规格里没写的改进。

    Awaiting translation

9/14Mon
  1. Hacker News · AI 编程经验问答20

    有人把遗留代码库改造得更易被 AI 编程智能体理解吗?

    一位创业公司首位工程招聘在 Hacker News 发问:如何把遗留代码库改造得更易被 AI 编程智能体理解。他所在的 2 年历史初创公司代码库始于 CTO 用 Lovable 搭的 MVP,如今存在业务逻辑重复、僵尸表和列、隐式依赖等问题。他计划把整个代码库文档化为 AI 智能体的“知识数据库”,并参考 Meta 的一篇文章,希望听取有类似经历的工程师分享做法、结果与避坑建议。

    Awaiting translation

9/12Sat
  1. Sean Goedecke · Blog62

    为什么不该专门为 AI 智能体打造工具

    Sean Goedecke 认为,大多数“为 AI 智能体打造 X”的尝试会失败,理由有三:适合智能体的工具同样适合人类,智能体像工程师一样输入文本、调用 API、读图和分类;现有工具已存在于训练数据中,新工具若只比人类版好 20% 也难被采用,这也是他不看好为智能体设计新编程语言的原因;智能体的理想工效学尚无定论,静态类型语言等说法都能讲出正反两套故事。

    Awaiting translation

9/11Fri
  1. OpenAI Developer Blog · Codex66

    OpenAI on How to Rewrite Skills and Prompts for GPT-6 Astra

    In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.

    Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.

9/10Thu
  1. Cursor · Changelog76

    Cursor launches “Projects,” a feature that uses a coordinating agent to take on long-running development work

    Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.

    Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.

  2. AI Coder · Telegram88

    Stolen Thoughts 研究:加密 reasoning block 可被跨模型解密,泄露 API key 与密码

    Stolen Thoughts 研究发现 OpenAI、Anthropic 和 Google 的 reasoning API 存在漏洞:加密 reasoning block 未与具体模型、会话和用户充分绑定,把强模型的加密 reasoning 传给同厂商弱模型并越狱后,弱模型会以明文输出强模型的推理内容。

    Awaiting translation

    Why it matters: 研究揭示加密 reasoning block 可跨模型解密,并给出公开轨迹中泄露密钥的实测数据,对智能体基础设施设计有直接参考价值。

9/9Wed
9/8Tue
9/6Sun
9/5Sat
  1. Ryan Lopopolo66

    An agent platform built for inventing agents: decoupling capability interfaces from their implementations

    Author Ryan Lopopolo argues that an agent is a parameterized program built on top of a set of capabilities: models and configurations, reasoning and tool-call loops, computers, disks, context, Skills, tools, connectors, runtimes, network policies, identity, IAM, guardrails, I/O channels, and system prompts.

    Why it matters: Drawing on his experience building multiple agents, the author proposes a platform architecture that decouples capability interfaces from their implementations — a useful reference for teams building Agent platforms.

  2. Vibe Code Textbook · Articles78

    如何写出编码智能体能完成的任务:六段式 spec 模板与 linter

    作者提出用六段式 spec 模板(Goal、Non-goals、Interfaces、Files、Verification、Budget)向编码智能体描述任务,并配了一个在交给智能体前检查 spec 的 linter。

    Awaiting translation

    Why it matters: 给出可直接套用的六段式 spec 模板、tally 实例和配套 linter,读者能据此改造自己交给编码智能体的任务描述。

  3. Vibe Code Textbook · Articles62

    解读 OpenHands 架构:step 循环、沙箱与 condenser 如何组织

    作者按 commit f7fb0c4b 阅读 OpenHands 仓库,梳理其架构:Agent Canvas 是托管编码智能体的控制中心,真正的智能体在 Software Agent SDK 中,采用单步执行模型,每步依次检查待确认动作、压缩事件历史、查询 LLM、解析响应、确认门禁、执行工具并生成 ObservationEvent。

    Awaiting translation

9/4Fri
  1. DevAgentStack · Field Notes82

    GPT-6 Astra vs. Fable 5.1 benchmarks: which scores are comparable and which aren't

    On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.

    Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.

9/3Thu
9/2Wed
  1. 陈与小金 · AI Coding 博客78

    Claude Code subagents: dispatch 13 AI workers at once, and the main conversation only gets 3 conclusions

    Drawing on a hands-on session where he dispatched 13 subagents to build a storyboard, the author walks through the subagents feature that both Claude Code and Codex have: subagents work in their own separate windows and hand only their conclusions back to the main conversation.

    Why it matters: Using a hands-on session where he dispatched 13 subagents to build a storyboard, the author shows how subagents keep their work outside the main conversation and send back only the conclusions.

  2. 宝玉78

    Anthropic's E-commerce AI Agent Engineering Guide: Architecture, Latency and Cost Optimization, and Production Practices

    Anthropic has published a guide dissecting e-commerce AI agents. Drawing on deployment experience with retailers, e-commerce platforms, and teams in travel, entertainment, and telecom, it proposes a single-agent architecture that puts Claude in a standard agent loop, uses skills to cover long-tail needs, and calls tools to work with existing systems. The guide says that in comparative testing, this architecture beats both sub-agent designs and the approach of cramming everything into the prompt.

    Why it matters: Drawing on enterprise e-commerce agent deployment experience, Anthropic lays out a complete engineering approach covering a single-agent-plus-skills architecture, latency and cost optimization, and memory and security evaluation.

  3. Hacker News · Agent Skills78

    mattpocock releases AI coding Agent Skills built for real engineering

    Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.

    Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.

8/31Mon
  1. Ryan Lopopolo38

    Harness Engineering 本质是即时上下文学习

    模型能力再强,其"好工作"的先验也未必与你的标准一致,因此始终需要围绕模型整理环境,让它在非功能性要求上偏向你或组织认可的选择。模型变好不会让这种上下文学习(ICL)需求消失。所谓 Harness Engineering,本质是一组技巧,为 ICL 提供即时机会,在不束缚推理模型的前提下对齐模型行为。

    Awaiting translation

  2. 宝玉71

    AI 原生思维:像训练大模型一样训练自己

    宝玉在演讲中提出 AI 原生思维,主张做 AI 产品要盯着模型能力边界线找需求,并按能力、成本、价值三条边界判断值不值得做。他以自己做的字幕翻译 App BaoCut 为例,说明从模拟字幕组的 V1 转向以终为始的 V2 后,用词级时间戳对齐、术语表注入和 Agent 自验证替代人工校对,一次成本优化把调用从 33 次降到 12 次、单集处理从 31 分钟降到 18 分钟。

    Awaiting translation

  3. Martin Alderson62

    用让智能体出题测验的方式减少代码库认知债

    作者提出用一段提示词让编码智能体就当前代码库出 5 道难度递增的题,答完后由它解释自己理解错的地方,提示词中要求使用 askuserquestiontool。作者认为这种方式比让智能体直接描述代码更有效,因为它掌握了自己所想与代码实际行为的差距;在复杂重构前先让智能体就方案出题,也能提前发现问题,同样的做法还可用于复杂表格或文档集,甚至可以要求通过 PR 测验后才允许智能体创建 PR。

    Awaiting translation