The UK AI Security Institute is publishing detailed frontier-model evaluation results through EvalEval, an open project that puts benchmark findings and their experimental context into a common format. The aim is to make expensive AI tests easier to verify, interpret and compare.
The first release covers five benchmarks: HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. It includes results for six Claude and GPT models, plus separate cyber evaluations. Configuration and transcript-level information is included where appropriate, rather than presenting only a final score.
That detail matters because benchmark performance can change with the evaluation protocol and the amount of computation allowed at test time. In Humanity’s Last Exam, for example, models continued solving more tasks as token use increased when they received correctness feedback after each attempt. A single leaderboard number can conceal those choices.
EvalEval uses its Every Eval Ever schema and Evaluation Cards to standardize reporting. The release will not make every result perfectly reproducible—rerunning frontier evaluations can remain costly—but it provides reference data for checking claims and studying how setup decisions shape reported performance.