Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
- Published
- Sep 29, 2026 — 17:33 UTC
Problem
This work addresses a gap in the estimation of confidence levels in large language models (LLMs) based on the probabilities of selected key tokens. Traditional methods often fail to accurately reflect the uncertainty in model predictions, particularly when assessing divergent outputs from multiple models. The authors propose a new approach to quantify this uncertainty, which is particularly relevant given the increasing reliance on LLMs in critical applications. Notably, this is a preprint and has not undergone peer review.
Method
The authors introduce the Divergent Token Confidence (DTC) framework, which quantifies reasoning uncertainty by counting tokens where two models exhibit disagreement. This is achieved through the application of Jensen-Shannon divergence, a method that measures the similarity between two probability distributions. The DTC framework supports both white-box and black-box evaluation types, allowing for flexibility in assessing model performance. Importantly, the method does not require explicit training, making it applicable across various model families. The experiments were conducted on six mathematical benchmarks, providing a robust evaluation of the proposed method's effectiveness.
Results
The DTC framework demonstrates significant improvements in Expected Calibration Error (ECE) metrics compared to standard full-sequence confidence methods. In white-box evaluations, the DTC achieved an ECE of 13.0%, markedly better than the 32.7%-42.4% range reported for traditional methods. In black-box evaluations on the DeepSeek-V3.2 benchmark, the DTC yielded an ECE between 13.7% and 16.3%, compared to the original verbalized scores that ranged from 32.1% to 40.2%. These results indicate a substantial enhancement in the calibration of model confidence estimates.
Limitations
The authors do not report any limitations in their study. However, the absence of explicit training may raise questions about the generalizability of the DTC framework across different tasks or domains, which is not addressed in the paper.
Why it matters
The implications of this work are significant for downstream applications of LLMs, particularly in scenarios where understanding model uncertainty is critical. By providing a more accurate measure of confidence through the DTC framework, this research could enhance the reliability of LLMs in decision-making processes, thereby fostering greater trust in AI systems. Furthermore, the methodology could inspire future research aimed at improving uncertainty quantification in various machine learning contexts.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
