Notabletraining methods

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

Xin Yu, Liuchen Liao, Yiwen Zhang, Yingchen Yu, Lingzhou Xue, Qinzhen Guo

Published
May 6, 2026 15:31 UTC

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI