Terminal-Bench ships new Harbor features, turning the benchmark into a versioned asset that keeps getting updated
The Terminal-Bench team has shipped a batch of new Harbor features that let datasets be released by version and let leaderboards migrate to new versions by reusing, re-evaluating, or rerunning trials. Tasks use semantic versioning: patch-level changes reuse old results as-is, validator changes only require re-evaluating saved artifacts, and only major changes that significantly alter the agent environment require a rerun. Dataset versions follow the highest version number among the tasks, and leaderboards use diffs to rerun only the tasks with major changes.
Why it matters: The Terminal-Bench team maintains the benchmark like software, laying out concrete mechanisms for task semantic versioning and leaderboard upgrades that you can carry over to your own evaluation pipeline.