Notableevaluation benchmarksHugging Face

AISI and EvalEval Release Reproducible Benchmark Results for AI Models

Published
Sep 22, 2026 00:00 UTC
Also in this story:UK AI Security Institute

On September 22, 2026, the UK AI Security Institute (AISI) published results from five benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench, and Terminal-Bench 2.0. These benchmarks evaluated six frontier models, including Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. Performance on the Humanity's Last Exam benchmark was noted to vary significantly based on the evaluation protocol and inference compute used. This initiative follows a joint workshop held in 2025 between AISI and the EvalEval Coalition, which aims to enhance the evaluation ecosystem through a shared reporting schema known as Every Eval Ever. AISI and EvalEval are collaborating to address gaps in evaluation reporting and to build a shared infrastructure that supports reproducibility in AI model assessments. The introduction of Evaluation Cards will facilitate the reporting of evaluation results, allowing researchers and practitioners to scrutinize individual studies more effectively.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: Hugging Face Blog