GitHub 与 Microsoft 开放 ReviewBench 评测 AI 代码评审
GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。
Awaiting translation
GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。
Awaiting translation
GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。
Awaiting translation
Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。
作者用 npm 发布时间戳统计 2026 年 Q3 的发布次数:Claude Code 发布 79 个版本、Codex CLI 38 个、Gemini CLI 15 个,中位发布间隔分别为 0.96 天、1.64 天和 6.09 天。
Awaiting translation
Linux Foundation 与 Agentic AI Foundation 于 9 月 14 日推出首个厂商中立的 MCP 认证 MCPA,考试为线上监考多选题,费用 US$250,有效期两年,含一次免费重考,对齐 2026-07-28 版 MCP 规范。
Awaiting translation
GPT-5.5 于 2026 年 4 月 23 日发布,是 OpenAI 自 GPT-4.5 以来首个完全重训练的基座模型,在 Terminal-Bench 2.0 上得分 82.7%,相同 Codex 任务下比 GPT-5.4 少用 40% token,但输入输出价格翻倍至每百万 token 5 美元和 30 美元。
Awaiting translation
9 月 15 日发布的 System One 决策模型 Jev 一周内 GitHub 相关项目超两千个、star 破四万,一批论文随之涌现。
Awaiting translation
作者上手体验了 OpenAI DevDay 发布的 Dots、Decisions API、开源 Codex Harness 更新和 ChatGPT Space 四项能力。
Awaiting translation
谷歌发布旗舰模型 Gemini 4 Argon,官方公布的 18 项测试中独占第一 12 项,DeepSWE v1.1 以 77.9% 位列全球第一,AutomationBench-AA 以 77.5% 领先 Claude Sonnet 5.5 的 71.3%。
Awaiting translation
俄罗斯团队 Маракуйя ИИ 的智能体框架首次参加 Terminal-Bench 4 和 OSWorld 2 两项基准测试,取得亮眼成绩。
Awaiting translation
On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.
Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.
Terminal-Bench 发布 4.0 版本,校准任务的时间、CPU 和内存资源,修复 19 个任务并移除 8 个饱和或存在质量问题的任务,所有任务统一设为 8 小时 agent 超时。
Awaiting translation
斯坦福大学研究人员主导、Terminal-Bench 团队联合全球科研机构专家打造的 Terminal-Bench-Science 0.1 发布,首批含生命、物理、地球、数学和工程科学领域的 70 项任务。
Awaiting translation
Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。
Awaiting translation
Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。
The Terminal-Bench team releases Terminal-Bench 3.0, whose first version spans 7 domains and 74 tasks, with the strongest model passing about 34%. Building on Terminal-Bench 2.1, this release broadens task diversity and adds CI/CD, semantic versioning, and result migration to keep improving the benchmark.
Why it matters: Terminal-Bench 3.0 rebuilds the benchmark with 74 tasks and CI/CD-based versioning, so readers can see how the new benchmark separates models.
The Terminal-Bench team has shipped a batch of new Harbor features that let datasets be released by version and let leaderboards migrate to new versions by reusing, re-evaluating, or rerunning trials. Tasks use semantic versioning: patch-level changes reuse old results as-is, validator changes only require re-evaluating saved artifacts, and only major changes that significantly alter the agent environment require a rerun. Dataset versions follow the highest version number among the tasks, and leaderboards use diffs to rerun only the tasks with major changes.
Why it matters: The Terminal-Bench team maintains the benchmark like software, laying out concrete mechanisms for task semantic versioning and leaderboard upgrades that you can carry over to your own evaluation pipeline.
Terminal-Bench 发布 Challenges,一种长周期、高 token 消耗的单任务基准,要求智能体在无时间与资源限制下自主完成整个项目,首批开放三个挑战。
Awaiting translation
Why it matters: Terminal-Bench 官方推出长周期单任务基准,并公开三个挑战的实测失败模式,可供评估智能体长时自主能力时参考。
Lovable 在早期访问中测试 GPT-5.5,其内部基准显示最难任务通过率从 GPT-5.4 的 36.9% 升至 41.6%,每次请求工具调用减少 23.1%,用户卡住的消息占比下降 9.9%。GPT-5.5 每请求输出 token 减少 33%,日常任务成本效率提升约 15%,将很快向 Lovable 构建者开放。
Awaiting translation
Terminal-Bench 团队宣布正在开发 Terminal-Bench-Science(TB-Science),面向生命科学、物理科学、地球科学及数学与计算科学等自然科学领域,目标构建 100+ 个可执行基准任务,在容器化环境中以确定性程序化验证评估 AI 智能体。
Awaiting translation
Terminal-Bench 3.0 已进入开发阶段,目标收录 100 个多样化任务,发布时最强模型的解决率不超过 30%。任务需为可通过命令行完成并程序化验证的真实计算机工作,覆盖更长周期、多微服务/文件系统/数据库等更丰富环境及专家级知识,合并窗口开放至 5 月底。贡献者提交一个被接受的任务即可在最终版本中获得署名。
Awaiting translation
Qodo 研究团队发布 AI 代码审查基准 1.0,向真实已合并 PR 注入缺陷,用 100 个 PR、580 个问题同时评测代码正确性与最佳实践合规性。在对比 Qodo 模型与 7 家主流 AI 代码审查平台的评测中,Qodo 以 60.1% 的 F1 分数领先,基准及评测结果已在 GitHub 公开。
Awaiting translation
Terminal-Bench has released version 2.0 and the Harbor package. The former is a more rigorously validated, harder benchmark for evaluating agents; the latter is for evaluating and optimizing agents. Harbor rewrites Terminal-Bench's test harness, supports deploying containers in the cloud, provides rollout interfaces for RL and SFT, and works with any agent you can put in a container.
Why it matters: Terminal-Bench 2.0 and Harbor are released together, so readers can see how the agent evaluation benchmark is validated and how to scale it in the cloud.
Terminal-Bench 发布数据集注册表,让智能体开发者通过 Terminal-Bench 框架统一评测多个智能体基准,基准开发者也能借此分发自己的基准。
Awaiting translation
Warp 在 Terminal-bench-core v0.1 上解决 52% 的任务,比此前最好成绩提升 9%,首次登顶该基准。此前 Anthropic 的 Opus 4 以 43% 的任务解决率创下纪录。
Awaiting translation
Terminal-Bench 新增 parallelize-graph 和 feal-linear-cryptanalysis 两个高难度任务,分别要求智能体用 Unified Parallel C 实现基因组组装的并行 de Bruijn 图构建,以及对 FEAL 类分组密码实施已知明文线性密码分析攻击。
Awaiting translation
Anthropic 将 Terminal-Bench 列为 Claude 4 模型卡七项基准之一,Claude 4 Opus 在 Terminal-Bench-Core 上取得 43.2% 的 SOTA 成绩。Dario Amodei 在 Code with Claude 主题演讲中也提及该基准。Terminal-Bench 团队表示将在未来几天验证 Claude 4 的表现并更新官方排行榜。
Awaiting translation
Terminal-Bench 发布首个版本,用于量化 AI 智能体在终端中执行复杂任务的能力,首发数据集 Terminal-Bench-Core-v0 包含 80 个手工编写并人工验证的任务,每个任务配有独立 Docker 环境、人工验证的解法与测试用例。
Awaiting translation
Why it matters: Terminal-Bench 给出 80 个带 Docker 环境和测试用例的终端任务,可用来横向比较不同智能体在命令行中的实际表现。