Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization
Huzi Cheng, Zhewei Zhang
- Published
- Sep 28, 2026 — 17:12 UTC
Problem
This work addresses a gap in understanding whether different forms of reasoning in language models, particularly in the context of multi-hop reasoning tasks, rely on the same underlying mechanisms. The authors explore this through the lens of depth generalization, specifically focusing on the performance of various model architectures. The paper is a preprint and has not undergone peer review.
Method
The authors propose five variants of the GPTNeoX architecture, each trained from scratch:
- Vanilla Model: The standard architecture without any modifications.
- Chain-of-Thought (CoT) Model: This variant incorporates reasoning steps directly into the output, allowing for a more explicit representation of the reasoning process.
- Pause Token Model: This model introduces tokens that indicate pauses in reasoning, potentially allowing for a more structured approach to multi-step reasoning.
- Latent-Reasoning Models: Two models are optimized end-to-end without intermediate reasoning traces, focusing on latent representations to capture reasoning dynamics.
The primary task evaluated is the extended multi-hop reasoning task known as ProsQA-Ext. The training compute used for these models is not specified in the paper.
Results
The results indicate that:
- In-Distribution (ID) Performance: The Vanilla, CoT, and Pause Token models demonstrate strong performance on in-distribution tasks.
- Out-of-Distribution (OOD) Generalization: The latent reasoning variants outperform the Vanilla, CoT, and Pause Token models in terms of OOD generalization, suggesting that these models are better at generalizing beyond the training distribution.
- Internal Dynamics: The internal dynamics of the latent reasoning models are consistent with forward reachability propagation on the graph, indicating a structured approach to reasoning.
- Causal Interventions and Circuit Analysis: The analysis reveals that computation is localized to a sparse recurrent search circuit within the bottleneck latent model, highlighting the efficiency of the latent reasoning approach.
Limitations
The authors note that strong performance on in-distribution tasks does not necessarily guarantee effective out-of-distribution generalization. Additionally, the Vanilla, CoT, and Pause Token models are found to rely heavily on local graph features, which may limit their generalization capabilities.
Why it matters
This research has significant implications for the design of language models, particularly in enhancing their reasoning capabilities. By demonstrating that latent reasoning can lead to improved OOD generalization, the findings suggest a potential shift in how future models are architected and trained, emphasizing the importance of underlying mechanisms in reasoning tasks.
By Callan Zhang · Sep 28, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
