Skip to content

#Reviews/benchmarks

0 items today
10/6Tue
  1. Tproger · Программирование58

    GitHub 与 Microsoft 开放 ReviewBench 评测 AI 代码评审

    GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。

    Awaiting translation

  2. GitHub Blog · Copilot66

    GitHub 发布 AI 代码评审开放基准 ReviewBench

    GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。

    Awaiting translation

    Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。

10/5Mon
10/4Sun
10/3Sat
10/1Thu
9/22Tue
9/4Fri
  1. DevAgentStack · Field Notes82

    GPT-6 Astra vs. Fable 5.1 benchmarks: which scores are comparable and which aren't

    On September 3, 2026, OpenAI released GPT-6 Astra and Astra Pro, initially limited to enterprises in the Daybreak cybersecurity program, with paid ChatGPT, the API, and AWS opening up over the following days.

    Why it matters: We break down the benchmark comparison between GPT-6 Astra and Fable 5.1, pointing out that the tested versions and harnesses differ across teams, so readers can judge which scores are actually comparable.

8/28Fri
8/27Thu
  1. Cline · Blog71

    Cline 实测八个模型做 IMO 2026:DeepSeek V4 Flash 以 0.12 美元拿到金牌线

    Cline 让八个模型在自家 harness 里做 IMO 2026 六道题,证明由 GPT-5.5 和 Claude Opus 5 双盲按 0–7 分制评分、Gemini 3.1 Pro 仲裁,金牌线为 29 分。

    Awaiting translation

    Why it matters: Cline 用同一套 harness 盲评八个模型做 IMO 2026,给出分数与单次成本对照,可看开源权重模型的实际性价比。

7/30Thu
  1. Terminal-Bench · News60

    Terminal-Bench 3.0 is out: 74 tasks across 7 domains, with the strongest model passing about 34%

    The Terminal-Bench team releases Terminal-Bench 3.0, whose first version spans 7 domains and 74 tasks, with the strongest model passing about 34%. Building on Terminal-Bench 2.1, this release broadens task diversity and adds CI/CD, semantic versioning, and result migration to keep improving the benchmark.

    Why it matters: Terminal-Bench 3.0 rebuilds the benchmark with 74 tasks and CI/CD-based versioning, so readers can see how the new benchmark separates models.

  2. Terminal-Bench · News62

    Terminal-Bench ships new Harbor features, turning the benchmark into a versioned asset that keeps getting updated

    The Terminal-Bench team has shipped a batch of new Harbor features that let datasets be released by version and let leaderboards migrate to new versions by reusing, re-evaluating, or rerunning trials. Tasks use semantic versioning: patch-level changes reuse old results as-is, validator changes only require re-evaluating saved artifacts, and only major changes that significantly alter the agent environment require a rerun. Dataset versions follow the highest version number among the tasks, and leaderboards use diffs to rerun only the tasks with major changes.

    Why it matters: The Terminal-Bench team maintains the benchmark like software, laying out concrete mechanisms for task semantic versioning and leaderboard upgrades that you can carry over to your own evaluation pipeline.

6/18Thu
  1. Terminal-Bench · News60

    Terminal-Bench 发布 Challenges 长周期智能体基准

    Terminal-Bench 发布 Challenges,一种长周期、高 token 消耗的单任务基准,要求智能体在无时间与资源限制下自主完成整个项目,首批开放三个挑战。

    Awaiting translation

    Why it matters: Terminal-Bench 官方推出长周期单任务基准,并公开三个挑战的实测失败模式,可供评估智能体长时自主能力时参考。

4/24Fri
  1. Lovable · Blog38

    Lovable 早期测试 GPT-5.5:最难任务通过率 41.6%,比 GPT-5.4 提升 12.5%

    Lovable 在早期访问中测试 GPT-5.5,其内部基准显示最难任务通过率从 GPT-5.4 的 36.9% 升至 41.6%,每次请求工具调用减少 23.1%,用户卡住的消息占比下降 9.9%。GPT-5.5 每请求输出 token 减少 33%,日常任务成本效率提升约 15%,将很快向 Lovable 构建者开放。

    Awaiting translation

3/8Sun
3/5Thu
  1. Terminal-Bench · News30

    Terminal-Bench 3.0 公开征集任务贡献

    Terminal-Bench 3.0 已进入开发阶段,目标收录 100 个多样化任务,发布时最强模型的解决率不超过 30%。任务需为可通过命令行完成并程序化验证的真实计算机工作,覆盖更长周期、多微服务/文件系统/数据库等更丰富环境及专家级知识,合并窗口开放至 5 月底。贡献者提交一个被接受的任务即可在最终版本中获得署名。

    Awaiting translation

2/5Thu
  1. Hacker News · AI Code Review 讨论42

    Qodo 发布 AI 代码审查基准 1.0:100 个 PR、580 个问题

    Qodo 研究团队发布 AI 代码审查基准 1.0,向真实已合并 PR 注入缺陷,用 100 个 PR、580 个问题同时评测代码正确性与最佳实践合规性。在对比 Qodo 模型与 7 家主流 AI 代码审查平台的评测中,Qodo 以 60.1% 的 F1 分数领先,基准及评测结果已在 GitHub 公开。

    Awaiting translation

11/7Fri
  1. Terminal-Bench · News62

    Terminal-Bench ships version 2.0 and an optimized Harbor evaluation package

    Terminal-Bench has released version 2.0 and the Harbor package. The former is a more rigorously validated, harder benchmark for evaluating agents; the latter is for evaluating and optimizing agents. Harbor rewrites Terminal-Bench's test harness, supports deploying containers in the cloud, provides rollout interfaces for RL and SFT, and works with any agent you can put in a container.

    Why it matters: Terminal-Bench 2.0 and Harbor are released together, so readers can see how the agent evaluation benchmark is validated and how to scale it in the cloud.

7/15Tue
6/25Wed
6/20Fri
5/23Fri
  1. Terminal-Bench · News32

    Anthropic 在 Claude 4 模型卡中引入 Terminal-Bench,Claude 4 Opus 创下 43.2% 新 SOTA

    Anthropic 将 Terminal-Bench 列为 Claude 4 模型卡七项基准之一,Claude 4 Opus 在 Terminal-Bench-Core 上取得 43.2% 的 SOTA 成绩。Dario Amodei 在 Code with Claude 主题演讲中也提及该基准。Terminal-Bench 团队表示将在未来几天验证 Claude 4 的表现并更新官方排行榜。

    Awaiting translation

5/19Mon
  1. Terminal-Bench · News62

    Terminal-Bench 发布首个终端智能体评测基准

    Terminal-Bench 发布首个版本,用于量化 AI 智能体在终端中执行复杂任务的能力,首发数据集 Terminal-Bench-Core-v0 包含 80 个手工编写并人工验证的任务,每个任务配有独立 Docker 环境、人工验证的解法与测试用例。

    Awaiting translation

    Why it matters: Terminal-Bench 给出 80 个带 Docker 环境和测试用例的终端任务,可用来横向比较不同智能体在命令行中的实际表现。