AISI and EvalEval Release Reproducible Benchmark Results for AI Models
- Published
- Sep 22, 2026 — 00:00 UTC
On September 22, 2026, the UK AI Security Institute (AISI) published results from five benchmarks: HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench, and Terminal-Bench 2.0. These benchmarks evaluated six frontier models, including Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. Performance on the Humanity's Last Exam benchmark was noted to vary significantly based on the evaluation protocol and inference compute used. This initiative follows a joint workshop held in 2025 between AISI and the EvalEval Coalition, which aims to enhance the evaluation ecosystem through a shared reporting schema known as Every Eval Ever. AISI and EvalEval are collaborating to address gaps in evaluation reporting and to build a shared infrastructure that supports reproducibility in AI model assessments. The introduction of Evaluation Cards will facilitate the reporting of evaluation results, allowing researchers and practitioners to scrutinize individual studies more effectively.
By Callan Zhang · Sep 22, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: Hugging Face Blog
