Skip to content

#Reviews/benchmarks

0 items today
8/19Wed
  1. Cline · Blog71

    Cline's Evaluation Methodology and Trace for Open-Weight Models

    Cline has open-sourced its evaluation methodology for open-weight models, along with a hill-climbing score and trace worth over one thousand dollars, available for download and analysis. The post lays out five heuristics from the Hill Climber's Checklist: set a North Star metric, quantify noise, break down failure modes by task/model/vendor, don't assume more thinking is always better, and keep a private evaluation set.

    Why it matters: Cline shares its evaluation methodology and a trace worth over a thousand dollars, offering five transferable hill-climbing heuristics that teams building their own harness can reference.

7/30Thu
  1. Terminal-Bench · News60

    Terminal-Bench 3.0 is out: 74 tasks across 7 domains, with the strongest model passing about 34%

    The Terminal-Bench team releases Terminal-Bench 3.0, whose first version spans 7 domains and 74 tasks, with the strongest model passing about 34%. Building on Terminal-Bench 2.1, this release broadens task diversity and adds CI/CD, semantic versioning, and result migration to keep improving the benchmark.

    Why it matters: Terminal-Bench 3.0 rebuilds the benchmark with 74 tasks and CI/CD-based versioning, so readers can see how the new benchmark separates models.

  2. Terminal-Bench · News62

    Terminal-Bench ships new Harbor features, turning the benchmark into a versioned asset that keeps getting updated

    The Terminal-Bench team has shipped a batch of new Harbor features that let datasets be released by version and let leaderboards migrate to new versions by reusing, re-evaluating, or rerunning trials. Tasks use semantic versioning: patch-level changes reuse old results as-is, validator changes only require re-evaluating saved artifacts, and only major changes that significantly alter the agent environment require a rerun. Dataset versions follow the highest version number among the tasks, and leaderboards use diffs to rerun only the tasks with major changes.

    Why it matters: The Terminal-Bench team maintains the benchmark like software, laying out concrete mechanisms for task semantic versioning and leaderboard upgrades that you can carry over to your own evaluation pipeline.

4/22Wed
  1. Augment Code · Blog88

    Augment Code Tests AGENTS.md: A Good File Is Like a Model Upgrade, a Bad One Is Worse Than Nothing

    Augment Code pulled dozens of AGENTS.md files from its own monorepo and used its internal benchmark suite AuggieBench to compare how the same tasks performed with and without the file. The best files delivered a quality boost equivalent to upgrading from Haiku to Opus, while the worst made the output worse than having no AGENTS.md at all.

    Why it matters: Augment Code used internal benchmarks to quantify how much the different ways of writing AGENTS.md actually differ, so readers can adjust their own repo's documentation structure accordingly.

11/7Fri
  1. Terminal-Bench · News62

    Terminal-Bench ships version 2.0 and an optimized Harbor evaluation package

    Terminal-Bench has released version 2.0 and the Harbor package. The former is a more rigorously validated, harder benchmark for evaluating agents; the latter is for evaluating and optimizing agents. Harbor rewrites Terminal-Bench's test harness, supports deploying containers in the cloud, provides rollout interfaces for RL and SFT, and works with any agent you can put in a container.

    Why it matters: Terminal-Bench 2.0 and Harbor are released together, so readers can see how the agent evaluation benchmark is validated and how to scale it in the cloud.