Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces
Ratish Puduppully, Pranabendu Misra, Paarth Iyer, Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati
- Published
- Sep 29, 2026 — 17:46 UTC
Problem
This work addresses a gap in the capability of AI models to verify natural-language thinking traces, particularly in the context of grade-school mathematics. The authors highlight that existing models often produce correct answers while generating invalid reasoning traces, which raises concerns about the reliability of these models in reasoning tasks. The paper is a preprint and has not undergone peer review.
Method
The authors utilize the iGSM benchmark, a synthetic dataset designed for grade-school mathematics, to evaluate the performance of models based on their reasoning traces. The evaluation focuses on two primary metrics: answer correctness and trace validity. The findings reveal that 31.6% of correct answers are associated with invalid traces. Notably, over half of the evaluated traces pass syntactic and arithmetic checks but fail semantic dependency checks, indicating a disconnect between surface-level correctness and deeper reasoning validity. The authors also explore the impact of training interventions, noting that non-minimal training traces lead to non-minimal outputs. Additionally, they investigate token shuffling, where shuffling tokens in 10% of training trace sentences maintains accuracy despite none of the traces passing verification.
Results
The key results indicate that 31.6% of correct answers are linked to invalid traces, highlighting a significant issue in trace validity. Furthermore, the accuracy achieved with token shuffling remains near-clean, suggesting that models can retain performance even when the reasoning traces are not valid. However, the available text does not report quantitative results beyond these findings.
Limitations
The authors acknowledge that the validity of traces is decoupled from answer correctness, particularly when considering out-of-distribution scenarios. This raises questions about the generalizability of their findings. Additionally, the presence of non-minimal outputs undermines the use of minimality as evidence of selective planning, suggesting that further investigation is needed to understand the implications of trace generation on model performance.
Why it matters
The implications of this research are significant for downstream work in AI reasoning and interpretability. By revealing the discrepancies between answer correctness and trace validity, the findings prompt a reevaluation of how models are trained and assessed in reasoning tasks. This work encourages the development of more robust verification methods for reasoning traces, which could enhance the reliability of AI systems in critical applications.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
