微软与 UCSB 提出 ScholarEvolve:用论文进化 Agent Harness
微软与加州大学圣巴巴拉分校研究者公开论文 ScholarEvolve,从已发表的 Agent 研究中寻找改进思路,写进 Harness 后用真实任务检验,执行任务的模型保持不变。
Awaiting translation
微软与加州大学圣巴巴拉分校研究者公开论文 ScholarEvolve,从已发表的 Agent 研究中寻找改进思路,写进 Harness 后用真实任务检验,执行任务的模型保持不变。
Awaiting translation
逛逛GitHub 盘点了 9 月份 GitHub 上热度最高的 20 个开源项目。Archify 以约 3.34 万新增 Star 居首,它能把系统描述或代码仓库画成交互式架构图;Ponytail 和 God's Eye View 分别以约 3.12 万、3.11 万 Star 紧随其后。
Awaiting translation
作者统计了公开 GitHub 仓库中 557 个 Cursor 规则文件,其中 510 个是 .cursor/rules 下的 .mdc 文件,47 个是旧的 .cursorrules 文件。
Awaiting translation
作者指出 .claudeignore 对 Claude Code 无效,官方文档要求改用 Read deny 规则,可在 .claude/settings.json 中拒绝 Read(.env)、Read(.env.*) 和 Read(./secrets/**),并用 /status 确认加载。
Awaiting translation
The author has GPT (gpt-6.1-sol, reasoning tier high) handle planning, key decisions, and acceptance in Codex, and calls DeepSeek-V4.1-Flash through the official DeepSeek Harness to write code, run experiments, and fix bugs—together producing a local PDF toolkit with five working features.
Why it matters: The author splits the work between GPT for planning and DeepSeek for execution to get a PDF toolkit running, and shares three prompts plus cache usage data that can carry over to cutting costs on long tasks.
作者从 Andrej Karpathy 一条讲如何看懂 AI 输出的 X 出发,实测了用具体风格名控制模型输出的做法。Karpathy 给出四个办法:用 ASD-STE100 写解释、让 AI 画图、直接输出 HTML 交互网页、为任意问题现做讲解视频。
Awaiting translation
On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.
Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.
作者因厌倦粘贴截图和用文字描述“左边第三个按钮”,开发了 Pinpoint,一个本地 MCP server 加 Chrome 扩展。在开发站点点击元素并输入要改的内容后,Cursor 会收到评论以及选择器、DOM 路径、计算样式、React/Vue 组件链、源文件提示和该元素的裁剪截图。
Awaiting translation
Nexusyn 发布开源长期记忆引擎,用 Go 编写,基于 PostgreSQL 18 与 pgvector,采用双时态版本化(valid_from / valid_to)让新事实自动取代旧事实并保留历史。
Awaiting translation
Rodion Mostovoy 与 Maxim Klyuchnikov 分享了在百万行级企业遗留项目上提升 AI 智能体效率的经验,包括构建术语表、处理技术债,以及用基准测试对比向量 RAG 与 GraphRAG 等上下文引擎的效果。
Awaiting translation
Andrej Karpathy 时隔两个月发文,讲述自己理解和消化 LLM 输出的方法,核心主张是随着 LLM 变强,人的工作会沿抽象层上移,变成监督和理解。
Awaiting translation
Karpathy 认为随着大语言模型变强,人的工作会更多上升到监督和理解层面,他分享了几个让模型输出更易读的技巧。
Awaiting translation
Matryoshka 是一个基于递归语言模型(RLM)思路的文档分析工具,让 LLM 输出受约束的 Nucleus S-expression 命令、由 Lattice 引擎执行,从而处理比上下文窗口大 100 倍的文件,无需向量数据库或分块启发式。
Awaiting translation
两位曾在多家 BI 公司工作的开发者开源了 Graphene,为编码智能体补上数据分析所需的上下文与语义层。
Awaiting translation
作者引用 Claude Code 团队在内部跑了几百个 skill 后的总结,提出 skill 不是一条保存下来的复杂提示词,而是一个文件夹,整个文件系统本身就是上下文工程。
Awaiting translation
开发者发布开源 Mac 应用 Deiko(MIT),通过本地 MCP 服务器让 Cursor 跨会话读取和更新每个任务的简要笔记,工具包括 search_briefs、get_task、list_tasks、save_outcome。
Awaiting translation
Personal Agent(个人智能体)被定义为持久记忆+工具执行+目标规划、以个体为中心的代理系统,是交互范式从"问—答"到"交代—办成"的迁移。其热度源于记忆基础设施成熟、MCP 标准化与模型同质化下的入口焦虑,工程重心在记忆质量、上下文工程与可靠性三处,开源侧可用 OpenClaw+Mem0+MCP 快速拼出原型。
Awaiting translation
作者认为 Personal Agent 是持久记忆加工具执行加目标规划的个体代理系统,交互范式从问答转向交代任务,本轮热度来自记忆基础设施成熟、MCP 标准化和巨头入口焦虑,而非单一技术突破。
Awaiting translation
Qoder 发布 v0.4.0,核心动作是把上下文从个人电脑搬进整个团队,新增项目(Projects)、讨论(Discussion)、自定义 Agent 和 Agent Team 四项能力,现已面向企业订阅开放。
Awaiting translation
一位开发者询问如何在 Codex 与 Claude Code 等 AI 编程智能体之间切换,以应对单个工具的用量限制,核心诉求是切换时不丢失正在进行的上下文和任务。他表示找到的少数方案都缺乏关注度,并说明自己目前无法承担从每个智能体每月 20 美元涨到 100-200 美元的套餐升级。
Awaiting translation
作者统计自己项目五周的 Claude Code 日志(112 个会话、584 次运行),发现被入口文件(.claude/commands 和 .claude/agents)点名的文档平均被打开 12.3 次,而没有任何入链的 88 个 markdown 文件中 57 个一次都没被打开,平均仅 1.2 次。
Awaiting translation
A product designer with six years of SaaS experience used Claude Code to single-handedly build Котомка, a life-planning app. Nearly all the code was written by AI; his job was to define requirements, review the results, and make decisions.
Why it matters: Using a real repository, the author documented the pitfalls he hit while building a product on Claude Code alone, plus the rules, hooks, and testing guardrails he set up around the AI.
作者根据 Claude 官方新出的提示词指南,整理出在 Claude Code 中使用 Opus 5.5 的几条做法。官方称 Opus 5.5 每次回复前都会自行决定思考多少,删掉提示词里的 think carefully 后回复更早且质量没有下降,作者用 grep 命令清理了 CLAUDE.md、rules 和 Skill 中的这类指令。
Awaiting translation
MemPO 在 GRPO 基础上为记忆片段额外计算 Memory Advantage,与轨迹级 outcome_adv 相加得到 final_adv,从而对"记忆写作"行为做细粒度引导。
Awaiting translation
国内首款个人AI助理 Today 中国版已正式上线,基础版永久免费,新注册用户可领 1 亿 token 和 1 个月 Pro 会员。作者基于 Today 搭建的个人助理“毛肚”接入本地文件夹、X、GitHub 和 Gmail 作为上下文,能主动提醒积压 17 天的 PR、按账号调性找选题并自动发布推文。
Awaiting translation
国内 Personal Agent 产品 Today AI 发布,可连接飞书、钉钉、Linear 等应用获取个人 Context,主动发现并推进任务,而非等待用户输入 Prompt,发布期可免费领取一亿 Token。作者认为它与 Meta Muse 同属 OpenClaw 思路的延续,差别在于把记忆做成产品层,让 Agent 从对话框转向主动协作。
Awaiting translation
文章拆解了 JEV 这类专注快速判断的决策模型在 Agent 系统中的 8 类应用场景,包括 ReAct 行动循环、Computer Use、工具与技能路由、工具安全守卫、任务评估与观测、模型与 Agent 路由、RAG 与上下文管理、实时动态决策。
Awaiting translation
MemorySync 发布面向 AI 智能体的托管记忆 API,在 Cursor 中以远程 MCP server 形式运行,可保存用户告知的决策、约定和约束,并在后续会话中检索,而不依赖聊天历史。
Awaiting translation
文章梳理了 LLM 与 Agent 系统的九类安全风险,包括 Prompt Injection、过度授权、身份与委托、敏感信息泄露、不安全的输出处理、供应链、RAG 与记忆投毒、成本放大和 agent 间通信不安全,并逐条给出防护做法。
Awaiting translation
面对 LLM 和 AI 智能体带来的行业剧变,资深工程师 Sean Goedecke 建议初级工程师不要轻信 ZIRP 时代的建议,不要参与政治斗争,保持友善和尽责。他同时提醒不要恐慌 AI,也不要完全回避 AI,而应把 AI 智能体的建议当作参考,用自己的判断去理解系统,避免成为只转发 AI 输出的“肉代理”。
Awaiting translation
Awaiting translation
Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels
Rust 开发者 Nikita 认为 AI 只加速了写代码这一段,上下文准备、审查、测试和修复仍要人来做,因此该按“到可上线改动”的时间算收益,而不是按模型首次响应算。
Awaiting translation
作者 fmind-dev 认为 AGENTS.md 与 Unix 工具链天然契合,因为配置、流和进程都能被智能体通过文件读写、运行检查并查看退出状态。
Awaiting translation
Lovable launches Chats, an agent that runs at the workspace level, can hold conversations across projects and trigger builds; once changes are confirmed, it hands the task off to the project's builder agent and brings progress back into the conversation.
Why it matters: Lovable shares the three-layer architecture behind Chats—trajectory, inbox, and activation—which you can adapt for your own multi-agent orchestration.
作者用 LangChain 的 MongoDBGraphStore 做对比测试,仅输入 5 份文档就自动生成 17 种实体类型、34 种关系类型,part_of、Part Of、part of 被识别为三类独立关联,导致检索时关联片段丢失。
Awaiting translation
TypeSafe 创始人 Diogo Almeida 发布笔记,提出把 coding agent 的 harness 围绕 Jev 重构:大模型负责写代码和复杂推理,Jev 负责选工具、挑上下文、路由请求等高频小决策,输入状态和问题,输出选项、评分或断言为真的概率,再由普通代码使用这些结果。
Awaiting translation
彼得·詹姆斯让 Meta 的 Muse 把可访问文件打包发到 Google Drive,结果收到 6.8 GB 的会话环境,包含内部文档、集成代码、记忆、日志和 SSH 密钥文件。
Awaiting translation
Cursor 官方指南给出三步配置,解决 Composer 在大型 monorepo 中跨包串上下文、混淆前后端约定的问题。做法是在各子包放置轻量 .cursorrules 而非单一根文件,因为 Composer 会优先采用离当前打开文件最近的规则文件;提示时显式引用相对路径(如 @packages/api/src/...),避免全工作区模糊语义搜索,让提示词预算集中在当前包的 AST 上。
Awaiting translation
Drawing on his own Codex setup, the author built a role-based model routing system for Claude Code: the main thread acts as coordinator, explorer uses Haiku for code search only, worker uses Opus for TDD implementation, verifier uses Sonnet to run checks independently, senior uses high-tier Opus for money, data, and concurrency, and reviewer switches to a different model for semantic review.
Why it matters: Based on a week of hands-on testing, the author shares the configuration, Hook enforcement, and cost trade-offs of multi-model division of labor in Claude Code, which you can adapt to your own multi-agent workflows.
作者指出负向提示词会把「不要」后面的词以高权重送进模型注意力,导致模型反而生成被禁止的内容,就像让人别想粉色大象。解法是剪枝优化:把「不要写长句子」这类黑名单改成正向闭集白名单,用具体动作替代抽象禁止,并把已弃用工具的历史叙事从规则文档中彻底删除。
Awaiting translation