Notableinterpretability

Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?

Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang, Luke Zhang, Haolin Ye, Xujie Si

Published
Sep 28, 2026 — 17:34 UTC

Problem

This work addresses a gap in the literature regarding circuit-based explanations, specifically their inadequacy in accounting for model errors. The authors highlight that existing methods fail to provide insights into why models, such as GPT-2, make mistakes, which is critical for improving interpretability and reliability in AI systems. The paper is a preprint and has not undergone peer review.

Method

The authors utilize the GPT-2 small model to evaluate circuit-based explanations through a series of experiments. They employ several datasets, including IOI, Docstring, and the Mechanistic Interpretability Benchmark, across six model-task settings. The evaluation metrics focus on exact answer agreement, measuring both model successes and failures. The methodology includes testing various ablation settings and circuit sizes to validate the circuits. Additionally, an error recovery method is introduced, which involves restoring omitted attention heads to improve the model's performance on error cases.

Results

The results indicate a high success agreement of 97.3-99.5% on correct prompts when compared to the model's correct answers. However, the error agreement is significantly lower, ranging from 11.4-41.7% on the model's errors. The proposed error recovery method shows a substantial improvement, increasing error recovery from 14.2% to 75.1% on a held-out set. Notably, the correct agreement decreases by 0.41 percentage points after implementing the error recovery technique, suggesting a trade-off between understanding errors and maintaining correct predictions.

Limitations

The authors acknowledge that while circuits can preserve task success, they do not adequately explain the reasons behind model failures. This limitation suggests that further research is needed to enhance the interpretability of model errors beyond what circuit-based methods currently offer. Additionally, the study's reliance on a single model (GPT-2 small) may limit the generalizability of the findings to other architectures or larger models.

Why it matters

This research has significant implications for the field of mechanistic interpretability in AI. By highlighting the shortcomings of circuit-based explanations in understanding model errors, it paves the way for future work aimed at developing more robust interpretability frameworks. The proposed error recovery method could serve as a foundation for enhancing model performance and understanding, ultimately contributing to the development of more reliable AI systems.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI