Skip to content
Original
Terminal-Bench · News·· 05/23/2025AI score32

Anthropic 在 Claude 4 模型卡中引入 Terminal-Bench,Claude 4 Opus 创下 43.2% 新 SOTA

Original title: Terminal-Bench on the Claude 4 Model Card

The title and summary in the selected language are awaiting translation.

AI overview

Anthropic 将 Terminal-Bench 列为 Claude 4 模型卡七项基准之一,Claude 4 Opus 在 Terminal-Bench-Core 上取得 43.2% 的 SOTA 成绩。Dario Amodei 在 Code with Claude 主题演讲中也提及该基准。Terminal-Bench 团队表示将在未来几天验证 Claude 4 的表现并更新官方排行榜。

Full text

Terminal-Bench on the Claude 4 Model Card

Anthropic features Terminal-Bench and sets a new SOTA

We're thrilled that Terminal-Bench was one of seven benchmarks featured on the Claude 4 model card and explicitely referenced as one of two benchmarks in their annoucement!

Terminal-Bench on the Claude 4 Model Card

Terminal-Bench was also one of two benchmarks mentioned by Dario Amodei during his keynote at the recent Code with Claude event:

Notably, while Claude 4 Opus' reported performance of 43.2% is state of the art on Terminal-Bench-Core, it is the lowest score on the model card. While other benchmarks are rapidly becoming saturated, Terminal-Bench is still a useful benchmark for evaluating the performance of frontier AI agents.

In the coming days we will be verifying Claude 4's performance on Terminal-Bench-Core and updating our offical leaderboard accordingly.

Source: Terminal-Bench · News · tbench.ai