GitHub 与 Microsoft 开放 ReviewBench 评测 AI 代码评审
GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。
Awaiting translation
GitHub 和 Microsoft 于 2026 年 10 月 5 日开放研究预览版 ReviewBench,用于评测 AI 代码评审智能体。该基准包含来自 187 个公开仓库、19 种语言的 219 个 pull request,语言和仓库规模分布基于对 GitHub 上 1.039 亿个 pull request 的分析,并刻意提高了实质性改动的占比。
Awaiting translation
GitHub 发布代码评审离线基准 ReviewBench,基于 1.039 亿个 GitHub PR 的分布特征,构建了覆盖 19 种语言、219 个公开 PR 的评测集,并公开数据集、评分规则与 LLM 评审模型配置。
Awaiting translation
Why it matters: GitHub 公开了 AI 代码评审基准的数据集、评分规则与评测入口,读者可据此对比不同评审智能体。
GitHub 发布 GitHub Agentic Workflows 0.89.22 预发布版,隔离任务不再使用 Docker sbx 和 gVisor,改为仅通过 Cloud Hypervisor 运行。
Awaiting translation
GitHub 发布 Spec Kit 1.0.11,目录添加命令对所有目录族改为幂等,并修复了 CommandRegistrar 中项目相对路径被重复转换的问题。
Awaiting translation
GitHub 于 9 月 23 日为 GitHub Copilot App 推出本地沙箱,目前处于公开预览,仅适用于使用本地仓库和工作树的会话。沙箱可分别限制文件系统读写路径、外网与局域网访问,以及 Git credentials 和 GitHub CLI 数据的使用,企业策略还能进一步收紧;若操作系统无法应用规则,隔离环境会报错退出而不执行命令。
Awaiting translation
GitHub has launched Project HydraFusion as a research preview in the Copilot CLI. It uses runtime orchestration to pick an execution plan across models from multiple providers. Users select it just like any other model, and billing follows each model's standard rates.
Why it matters: GitHub lays out three orchestration modes for HydraFusion and compares cost versus quality across three benchmarks, so you can judge the trade-offs of multi-model orchestration on real coding tasks.
Vercel 发布新版 v0,将其定位从生成演示转向生产级应用和智能体。新版本基于沙箱运行时,可导入任意 GitHub 仓库并自动拉取 Vercel 上的环境变量和配置;新增 Git 面板,让非工程成员也能为每个对话建分支、向 main 提 PR 并在合并后部署;同时提供与 Snowflake 和 AWS 数据库的安全集成,以及默认开启的部署保护和访问控制。
Awaiting translation
Why it matters: v0 从生成演示转向生产级应用,给出导入 GitHub 仓库、Git 面板和数据库集成等具体能力变化。