Notablereasoning

Base Models Can Reason By Taking a Cue From Training Data

Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros

Published
Oct 5, 2026 — 17:59 UTC

Problem

This work addresses a gap in understanding how training data influences the reasoning behavior of base models. The authors explore the impact of specific cues derived from training data on model performance, which is particularly relevant given the increasing reliance on large pre-trained models in various applications. The paper is a preprint and has not undergone peer review.

Method

The authors employed two models, Olmo-3-7B and Qwen3-14B, to evaluate the influence of starting token cues on reasoning tasks. They fixed specific tokens to assess performance, particularly focusing on the MATH-500 benchmark, measuring pass@1 accuracy as the primary performance metric. Causal data interventions were utilized, where the authors manipulated words in prompts to evaluate the effectiveness of these cues. Two types of prompt instructions were compared: "Think duck duck goose" and "Think step by step." Additionally, the authors analyzed hidden state representations to correlate them with document types from the training data, providing insights into how these representations are influenced by the training corpus.

Results

The results indicate a significant improvement in reasoning performance when using specific cues. The Olmo-3-7B model achieved a MATH-500 pass@1 accuracy of 78% when using cues, compared to 42% without cues. Similarly, the Qwen3-14B model demonstrated an accuracy of 87% with cues, up from 72% without. The effectiveness of the prompt instructions was also notable, with "Think duck duck goose" performing comparably to "Think step by step," suggesting that the choice of prompt can significantly influence model reasoning outcomes.

Limitations

The authors did not report any limitations in their study. However, the absence of reported limitations may suggest a need for further exploration of potential confounding factors or the generalizability of the findings across different tasks and datasets.

Why it matters

This research has significant implications for the design and training of base models, particularly in understanding how to leverage training data to enhance reasoning capabilities. By demonstrating that specific cues can lead to substantial performance improvements, the findings encourage further investigation into the mechanisms by which training data shapes model behavior. This could inform future model architectures and training strategies, ultimately leading to more effective AI systems.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI