Notableother

Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav

Published
Sep 30, 2026 — 17:48 UTC

Problem

This work addresses a gap in the literature regarding the effectiveness of unlearning in multilingual models, specifically focusing on cross-lingual loopholes. The authors highlight that existing methods may not adequately remove knowledge from models across different languages, leading to unintended retention of information. This issue is particularly relevant in the context of language budgeted multilingual unlearning, where the goal is to selectively forget information while maintaining performance across multiple languages. The paper is a preprint and has not yet undergone peer review.

Method

The authors introduce an algorithm named COVER, designed to optimize the selection of source languages for unlearning tasks. The algorithm aims to maximize the predicted coverage of languages that do not receive direct forget supervision. The evaluation is conducted on a benchmark called the Cross-Lingual Unlearning Tensor, which encompasses 174 language-script pairs and utilizes 25 atomic paraphrase types. The data source for this study is the Low Resource Languages for Emergent Incidents (LORELEI) corpus. Specific details regarding the training compute used for the experiments are not disclosed.

Results

The proposed method demonstrates a mean held-out residual access reduction ranging from 7.8% to 27.3% when compared to a baseline that employs uniform source selection. This indicates that the COVER algorithm effectively enhances the unlearning process by strategically selecting source languages.

Limitations

The authors note that naively selecting strong individual sources does not necessarily lead to the formation of effective source sets for unlearning. This limitation suggests that while the COVER algorithm improves upon existing methods, there may still be challenges in achieving optimal performance across all scenarios. Additionally, the lack of specified training compute may hinder reproducibility and further analysis of the method's efficiency.

Why it matters

The implications of this work are significant for future research in multilingual model management and unlearning. By addressing the cross-lingual loopholes, the findings could lead to more robust unlearning techniques that ensure sensitive information is effectively removed from multilingual models. This has potential applications in privacy-preserving AI systems and could inform the development of more sophisticated language models that are capable of ethical knowledge management.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI