Notabletraining methods

Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu

Published
Sep 29, 2026 — 17:07 UTC

Problem

The paper addresses a gap in the existing literature on on-policy distillation (OPD) of large language models, specifically the limitation of vanilla OPD methods that treat all teacher signals equally. This approach fails to account for the varying impacts that different teacher signals can have on student performance. The authors propose a solution to this issue through a more nuanced weighting mechanism for teacher signals, which is critical for improving the effectiveness of the distillation process. Notably, this work is presented as a preprint and has not yet undergone peer review.

Method

The authors formulate a bilevel optimization problem to achieve weighted OPD. The core of their method involves an iterative solver that alternates between updating token weights and the student policy. The loss function employed is a weighted OPD objective, which allows for the differentiation of teacher signals based on their effectiveness. The paper does not specify the data used for training or the compute resources required for the experiments. Additionally, the method includes a mechanism for updating weights in closed form, followed by a single gradient step, which enhances the efficiency of the optimization process.

Results

The proposed method demonstrates a significant improvement in average math performance, achieving an increase of 9.7 points compared to the baseline of vanilla OPD. This result highlights the effectiveness of the weighted approach in optimizing student performance through more selective teacher signal utilization.

Limitations

The authors do not report any limitations in their work. However, the lack of specified data and training compute details may hinder reproducibility and practical application in real-world scenarios.

Why it matters

This research has important implications for the field of machine learning, particularly in the context of distilling large language models. By introducing a method that optimally weights teacher signals, the findings could lead to more efficient training processes and improved performance in downstream tasks. This work paves the way for further exploration into adaptive distillation techniques that can leverage the strengths of various teacher models, potentially enhancing the capabilities of language models in diverse applications.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI