Notabletraining methods

Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution

Erfan Baghaei Potraghloo, Seyedarmin Azizi, Arya Fayyazi, Saeid Shokoufa, Mehdi Kamal, Souvik Kundu, Massoud Pedram

Published
Oct 5, 2026 — 17:53 UTC

Problem

This work addresses the challenge of language models producing incorrect answers despite higher probabilities assigned to correct responses. The authors highlight that existing methods often fail to leverage the full potential of probability distributions during answer generation. This paper is a preprint and has not undergone peer review.

Method

The authors propose On-Policy Power Distillation (OPPD), a novel approach that utilizes a sequential Monte Carlo sampler to train a model for generating answers in a single generation. The method employs a frozen teacher model that provides power distribution weights for candidate answers. During training, a maximum-likelihood update is applied, where each answer is weighted uniformly by the probabilities derived from the teacher model. The data sources used for training include MATH500, GSM8K, and HumanEval. The paper does not specify the training compute resources utilized.

Results

The results demonstrate significant improvements in accuracy across various benchmarks:

  • MATH500: OPPD achieves a 23.0-point increase over an untrained model.
  • GSM8K: OPPD shows a 27.3-point increase over an untrained model.
  • Single generation score: OPPD improves by 2.4 points over published power sampling with 64 candidates.
  • GRPO comparison on MATH500: OPPD scores 3.8 points higher than GRPO.
  • GRPO comparison on GSM8K: OPPD scores 4.0 points higher than GRPO.
  • GRPO comparison on AIME: OPPD scores 5.4 points higher than GRPO.
  • Cumulative score with GRPO: OPPD adds up to 9.3 points.
  • HumanEval: OPPD achieves an accuracy increase of up to 5.3 points.
  • The sharpening exponent is reported to range between 1.19 and 2.02.

Limitations

The authors do not report any limitations in their work. However, the lack of specified training compute resources may hinder reproducibility and scalability assessments.

Why it matters

The implications of this research are significant for downstream applications in natural language processing, particularly in enhancing the reliability of language models in generating accurate responses. By effectively utilizing power distributions, OPPD could lead to more robust models that better align with human-like reasoning, thereby improving user trust and applicability in critical domains.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI