Notableevaluation benchmarks

MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

Negin Kafee Hernashki, Soumick Chatterjee

Published
Oct 1, 2026 — 17:44 UTC

Problem

Unsupervised anomaly detection (UAD) methods for brain MRI currently lack a standardized evaluation framework, leading to rankings influenced by unreported methodological choices. This paper presents MIRTO, a registration-gated evaluation protocol designed to provide a more robust assessment of UAD methods in this domain. The work is a preprint and has not undergone peer review.

Method

MIRTO employs a comprehensive evaluation protocol that includes:

  • Registration Check: Ensures that the geometry of comparisons is gated with a registration check to maintain spatial consistency.
  • Threshold Setting: Establishes thresholds based solely on validation data, minimizing bias from test data.
  • False-Positive Volume Reporting: Quantifies the false-positive volume realized on the test set, providing a clearer picture of model performance.
  • Evaluation Pipelines: Utilizes 15,552 defensible evaluation pipelines to ensure robustness in results.
  • Bootstrap Intervals: Implements paired subject-bootstrap intervals with multiplicity control to enhance statistical validity.
  • Data Source: Evaluates four UAD methods trained on healthy data and tested on a dataset comprising 312 subjects from the BraTS 2020 challenge.

Results

The evaluation of UAD methods using MIRTO yielded the following results:

  • Voxel AUROC: Ranged from 0.873 to 0.583, significantly affected by axis-order mismatch, compared to a baseline diffusion model.
  • Variance Explained: Achieved at least 0.95 in voxel AUROC and area under the precision-recall curve (AUPRC), 0.77 in Dice coefficient, and 0.14 in lesion sensitivity.
  • Dice Advantage: Demonstrated significant improvements at validation thresholds; however, this advantage diminished when equalizing the realized false-positive burden.
  • REFLECT's Latent Aggregation: Enhanced the Dice score by 0.052 when controlling for equal burden, indicating potential for improved segmentation accuracy.

Limitations

The authors note that all inference drawn from the study is exploratory, as the same cohort was utilized for both protocol development and evaluation. This could introduce biases and limit the generalizability of the findings. Additionally, the reliance on validation data for threshold setting may not reflect real-world performance.

Why it matters

The introduction of MIRTO has significant implications for the field of unsupervised anomaly detection in medical imaging. By providing a structured evaluation framework, it allows for more reliable comparisons between different UAD methods, potentially leading to improved diagnostic tools in clinical settings. The protocol's emphasis on registration and statistical rigor may also inspire further research into standardized evaluation practices across various domains of medical imaging.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI