Notableinterpretability

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España, Jose M. Juarez, Juan Moreno-Garcia

Published
Oct 1, 2026 — 17:34 UTC

Problem

This work addresses the lack of consensus in Explainable Artificial Intelligence (XAI) due to the divergent explanations produced by various attribution methods. The authors highlight the need for a systematic approach to evaluate and compare the explanations generated by different techniques in the context of medical abstract classification. The study is presented as a preprint, indicating that it has not yet undergone peer review.

Method

The authors utilize the DeBERTa-v3 architecture for their experiments. The dataset comprises a Medical Abstracts corpus, with one thousand texts per class, allowing for a robust evaluation across multiple diagnostic categories. They formulate five enriched hypotheses for each diagnostic category to enhance the interpretability of the model's predictions. The explanation methods employed include SHAP, LIME, occlusion, Input x Gradient, and Attention x Gradient. To quantify the agreement between different explanation methods, the Jaccard index is used for pairwise comparison. Notably, the paper does not disclose the training compute resources utilized for model training.

Results

The predictive accuracy of the model is reported to be high across well-defined clinical domains, although specific numerical results or comparisons to baseline models are not provided. The authors observe performance degradation under conditions of high semantic ambiguity, indicating that the model struggles to maintain accuracy when faced with complex or unclear inputs. They also report strong explanatory stability in univalent categories, suggesting that the model's predictions are consistent and reliable in these cases. Additionally, the authors identify several error mechanisms, including lexical hypersensitivity, semantic overlap, and loss of attribution coherence, which contribute to the model's performance challenges. However, no quantitative results are provided for these observations.

Limitations

The authors acknowledge that the model's performance degrades significantly under high semantic ambiguity, which poses a challenge for practical applications in medical contexts. They also identify systemic failure mechanisms that could impact the reliability of the model's predictions and explanations. These limitations highlight the need for further research to enhance the robustness of explainability methods in ambiguous scenarios.

Why it matters

This work has significant implications for the development of explainable AI systems in healthcare, particularly in the context of zero-shot learning. By providing a comparative framework for evaluating different explanation methods, the authors contribute to the understanding of how various techniques can be leveraged to improve interpretability in clinical applications. The identification of error mechanisms also paves the way for future research aimed at mitigating these issues, ultimately enhancing the trustworthiness and usability of AI systems in medical decision-making.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI