Notable interpretability

Attention is Case-Sensitive

Maximilian Dillitzer, Tin Stribor Sohn, Jason J. Corso, Michael Auerbach

Published
Aug 4, 2026 — 14:16 UTC

Problem — This work addresses the under-explored impact of letter casing on attention allocation in Large Language Models (LLMs) and Vision-Language Models (VLMs), presenting a systematic empirical characterization. The study is a preprint and unreviewed, contributing to the understanding of how typographic features affect model behavior.

Method — The authors analyze 13 models, including nine LLMs and four VLMs, employing various tokenization schemes. They investigate how formatting target information in uppercase or alternating case influences internal attention allocation. The study frames the casing effect as a latent property of pretrained transformers rather than a prescriptive method. The authors also explore the interaction between attention and performance, noting that increased attention concentration does not necessarily correlate with improved task accuracy, particularly in high-entropy contexts.

Results — The available text does not report quantitative results.

Limitations — The authors note that while the casing effect robustly shifts attention, its impact on downstream accuracy is complex and can lead to performance degradation in certain contexts. They also identify that reasoning models exhibit a

Turing Wire

By Callan Zhang · Aug 4, 2026 · Editorial standards →

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.CL