Majoralignment safety

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin

Published
Sep 23, 2026 16:13 UTC

Problem

The paper addresses a significant gap in the capability for transparent and traceable empirical evidence in AI safety assessments, particularly in the context of the EU AI Act's Code of Practice. The authors highlight the need for a systematic approach to evaluate systemic risks associated with AI systems, which is currently lacking in the literature. This work is presented as a preprint and has not yet undergone peer review.

Method

The authors propose a Systemic Risk Index, which is an open evaluation pipeline and dashboard designed to facilitate the assessment of systemic risks in AI. The evaluation framework incorporates 19 public benchmarks categorized into four domains: Chemical, Biological, Radiological, and Nuclear (CBRN), cyber offense, harmful manipulation, and loss of control. The evaluation method employs harm-preserving perturbations and simulated deployment contexts to assess the risks associated with AI systems.

The dashboard features allow users to toggle between average and worst-case aggregation of risk scores, enabling a nuanced understanding of model capabilities and their impact on aggregate risk scores. Additionally, the dashboard provides traceability of risk ratings back to the benchmark evidence, enhancing transparency in the evaluation process.

Results

The results indicate a notable score reduction of 14 to 37 points when using worst-case aggregation compared to average assessments, highlighting the potential variability in risk evaluation based on aggregation methods. An agreement metric of $κ= 0.78 ext{--}0.82$ was observed between LLM judges and human graders, suggesting a high level of consistency in evaluations. Furthermore, the transformation preservation rate was reported at $83 ext{%}$, indicating that a significant majority of sampled transformations maintain the original harm characteristics. A survey of 21 participants revealed that the reported scores were easy to understand and that the dashboard effectively encourages varied evaluations.

Limitations

The authors do not report any limitations in their work, which may suggest a need for further exploration of potential shortcomings or areas for improvement in the evaluation pipeline and dashboard.

Why it matters

This work has significant implications for downstream research and practice in AI safety assessments, particularly in regulatory contexts. By providing a structured and transparent framework for evaluating systemic risks, the proposed pipeline and dashboard can enhance the accountability and reliability of AI systems. This could lead to more informed decision-making by stakeholders and regulators, ultimately contributing to safer AI deployment in sensitive applications.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI