Dex Horthy 给 @humanlayer_dev 的上下文工程建议:会话在压缩或新开时丢失上下文,应把所有上下文和决策写进 artifacts 里的文档(设计。
Awaiting translation
Dex Horthy 给 @humanlayer_dev 的上下文工程建议:会话在压缩或新开时丢失上下文,应把所有上下文和决策写进 artifacts 里的文档(设计。
Awaiting translation
Karpathy 认为随着大语言模型变强,人的工作会更多上升到监督和理解层面,他分享了几个让模型输出更易读的技巧。
Awaiting translation
Awaiting translation
Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels
Lovable launches Chats, an agent that runs at the workspace level, can hold conversations across projects and trigger builds; once changes are confirmed, it hands the task off to the project's builder agent and brings progress back into the conversation.
Why it matters: Lovable shares the three-layer architecture behind Chats—trajectory, inbox, and activation—which you can adapt for your own multi-agent orchestration.
In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.
Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.
Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.
Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.
GitHub Copilot 团队复盘了四项降低 AI 编码成本的改动,核心原则是按完整任务而非单次工具调用衡量效率。
Awaiting translation
Why it matters: GitHub Copilot 团队复盘四项降本改动,并给出可迁移的评估方法:按完整任务而非单次工具调用衡量成本。
Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.
Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.
OpenAI 发布 Codex Remote 使用指南,介绍如何在 ChatGPT 移动端启动、指挥、审查和整理运行在开发机上的编码任务,核心思路是把手机当作控制平面而非终端。
Awaiting translation
Why it matters: OpenAI 官方梳理 ChatGPT 移动端 Remote 控制 Codex 的完整用法,涵盖 Queue 与 Steer、side chat、Plan 与 Goal 等关键决策点。
AI Hero 的 skills 目录发布 v1,通过在各 Skill 上启用 disable-model-invocation: true,让 Skill 描述不再进入模型选择 Skill 时查看的上下文窗口,Skill 描述的 token 成本降低 63%。
Awaiting translation
Why it matters: v1 用 disable-model-invocation 把 Skill 描述移出上下文窗口,并区分用户调用与模型调用,读者可据此判断自己的 Skill 组织方式。
Augment wires the Incident Investigator expert from its internal Cosmos platform into Slack and PagerDuty. It automatically triages every alert and runs root-cause analysis, then suggests one of four actions: fix the code, roll back, upgrade, or just keep monitoring. Humans only review the RCA and make the call.
Why it matters: Augment has shared the full playbook for putting Cosmos Expert on alert triage, along with a month of before-and-after data, so you can adapt it to your own on-call process.
Lovable 团队为自家智能体搭建了两个自动化闭环,用来持续减少用户卡住的情况。第一个是 Lovable Stack Overflow(LSO)知识库,在用户请求前由分类器、选择器和合成器判断是否注入解决方案,早期版本让卡住率下降 5%、发布率提升 2%。
Awaiting translation
Why it matters: Lovable 团队公开两个自动化闭环的落地细节,可借鉴如何用知识库和反馈工具降低用户卡住率。
Anthropic 工程师 Thariq Shihipar 提出用 HTML 替代 Markdown 作为 Claude Code 的输出格式,理由是 HTML 信息密度更高、更易阅读和分享,还能做双向交互。
Awaiting translation
Why it matters: Anthropic 工程师分享用 HTML 替代 Markdown 承载 Claude Code 输出的做法,附常见场景的提示词与模板。
AI Hero 在其 skills 仓库新增 /handoff 和 /prototype 两个 Skill。
Awaiting translation
Why it matters: 作者公开了 /handoff 与 /prototype 两个 Skill 的设计思路,可看到上下文交接与原型验证如何嵌入智能体工作流。
Awaiting translation
@karpathy and I are back! At @sequoia AI Ascent 2026. And a lot has changed. Last year, he coined “vibe coding”. This year, he’s never felt more behind as a programmer. The big shift: vibe coding raised the floor. Agentic engineering raises the ceiling. We talk about what it means to build seriously in the agent era. Not just moving faster. Building new things, with new tools, while preserving the parts that still require human taste, judgment, and understanding.
AI Hero 的 skills 仓库更新,把 /ubiquitous-language 废弃并合并进新 Skill /grill-with-docs,输出从 ubiquitous-language.md 改为 context.md,并支持多个限界上下文各自维护共享语言。
Awaiting translation
Why it matters: 作者把 /ubiquitous-language 合并为 /grill-with-docs,并给出 ADR 触发条件与多限界上下文做法,可迁移到自己的 Skill 配置。
In AGENTS.md, Augment Code front-loads roughly 2.5k characters of Karpathy-style coding rules, then runs 40 OpenClaw PRs through Auggie, Claude Code, and Codex for comparison.
Why it matters: A head-to-head test of three coding agents on the same set of PRs shows that prompt constraints mainly cut costs rather than improve quality, and it also surfaces differences between the harnesses.
DX 对 500 家公司的纵向研究显示,AI 带来的 PR 速度中位提升为 7.5%,平均 13%,最高 70%。DX CTO Justin Reock 指出,工程师只有约 16% 的时间在写代码,只优化这一环,个位数提升就是预期结果;真正决定产出的是系统,包括代码模块化、文档、CI/CD 速度与智能体编排。
Awaiting translation
Augment Code pulled dozens of AGENTS.md files from its own monorepo and used its internal benchmark suite AuggieBench to compare how the same tasks performed with and without the file. The best files delivered a quality boost equivalent to upgrading from Haiku to Opus, while the worst made the output worse than having no AGENTS.md at all.
Why it matters: Augment Code used internal benchmarks to quantify how much the different ways of writing AGENTS.md actually differ, so readers can adjust their own repo's documentation structure accordingly.
OpenAI 用 GPT-5.3-Codex 在 Extra High 推理档下从空仓库连续运行约 25 小时、消耗约 13M token、生成约 3 万行代码,做出一个可测试的设计工具。
Awaiting translation
Why it matters: 作者用 25 小时、13M token 的实测展示长时程智能体如何靠持久化项目记忆和逐里程碑验证保持不跑偏。
Dagster Labs 分享了用 OpenAI Codex 加速技术文档写作、跨媒介内容转换和文档覆盖度评估的实践。他们重写了 CONTRIBUTING.md,明确文档层级、结构和最佳实践,让 Codex 能据此生成符合规范的文档;还借助 gh 命令让 Codex 解读 PR 的 diff 和描述,并让 Codex 把教程改写成 YouTube 视频脚本。
Awaiting translation
Why it matters: Dagster 团队把 Codex 用于文档写作、PR 解读和内容跨媒介转换,其中用文档生成代码来反向衡量文档覆盖度的做法可以迁移。