微软与 UCSB 提出 ScholarEvolve:用论文进化 Agent Harness
微软与加州大学圣巴巴拉分校研究者公开论文 ScholarEvolve,从已发表的 Agent 研究中寻找改进思路,写进 Harness 后用真实任务检验,执行任务的模型保持不变。
Awaiting translation
微软与加州大学圣巴巴拉分校研究者公开论文 ScholarEvolve,从已发表的 Agent 研究中寻找改进思路,写进 Harness 后用真实任务检验,执行任务的模型保持不变。
Awaiting translation
On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.
Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.
Awaiting translation
Claude Tag in Slack can now use your personal connectors! You can now securely access that Drive doc, Salesforce account, or Warehouse table that you have personal access to right where the work happens. Avail. on Teams today and Enterprise next week https://claude.com/blog/claude-tag-now-supports-personal-connectors-in-channels
Cursor introduces “Projects,” a feature built for long-running work like a single feature, a migration, or an entire application. It keeps context over months and delegates tasks to thousands of sub-agents. Projects are powered by cloud agents: the coordinating agent doesn’t write code, it only plans, assigns work, and hands back results, spinning up local agents when on-device testing is needed. Each project keeps a set of files synced between the cloud and local machines, steadily accumulating research findings, artifacts, and knowledge of the codebase.
Why it matters: The official docs lay out the context-sharing and auto-triggering mechanisms for project-based multi-agent collaboration, which you can use to judge how long-running tasks get taken over.
On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.
Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.
Cursor has updated its cloud agents and Cursor harness so cloud agents can subscribe to event sources, resume when there's new activity in a PR, Slack thread, or scheduled task, and keep going until the work is done—fixing CI failures and handling bot comments.
Why it matters: Cloud agents are moving from one-shot runs to subscribing to events and following up continuously on PRs and Slack threads, which gives readers a way to judge how the boundaries of automation are shifting.
Codex 团队在 GPT-5.6 发布后于 Reddit 举办 AMA,说明 Sol 是主力模型、Terra 更快更省、Luna 主要用于廉价子智能体和上下文收集,UI 场景推荐用 Sol 配合参考图。
Awaiting translation
Anthropic 为 Opus 4.6 和 Sonnet 4.6 推出 1M 上下文窗口正式版,作者实测约 500K token 的 Claude Code 会话仍能保持任务连贯。
Awaiting translation