Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
- Published
- Oct 1, 2026 — 17:45 UTC
Problem
This work addresses the limitations of supervised finetuning (SFT) and reinforcement learning (RL) in posttraining, particularly focusing on the challenges of generalization and catastrophic forgetting. The authors highlight that existing methods often struggle to maintain performance across diverse tasks and data distributions, which is critical for deploying models in real-world applications. This paper is a preprint and has not undergone peer review.
Method
The authors propose a novel algorithm that employs a Markov chain Monte Carlo (MCMC) sampling technique to enhance the finetuning process. This method transforms off-policy traces to be more on-policy by utilizing a reference model, thereby improving the quality of the training data used during finetuning. The approach aims to leverage the strengths of SFT while mitigating its weaknesses by ensuring that the samples drawn during training are more representative of the target distribution.
Results
The results indicate that the proposed SFT method rivals existing posttraining techniques, particularly against strong on-policy baselines. Specifically, the finetuned models demonstrate superior generalization performance, effectively maintaining performance across various tasks. Additionally, the authors report that SFT exhibits less catastrophic forgetting compared to these strong on-policy baselines, suggesting a more robust retention of learned information. However, the available text does not report quantitative results for distributional performance or specific metrics against the baselines.
Limitations
The authors do not report any limitations in their study. However, it is important to note that the lack of quantitative results may limit the ability to fully assess the performance of the proposed method against established benchmarks. Furthermore, the reliance on a reference model for data transformation could introduce biases depending on the quality and representativeness of the reference model used.
Why it matters
This research has significant implications for downstream work in the field of machine learning, particularly in the context of model deployment in dynamic environments where data distributions may shift. By demonstrating that SFT can achieve competitive performance while reducing catastrophic forgetting, this work opens avenues for further exploration of finetuning techniques that can adaptively learn from new data without losing previously acquired knowledge. This could lead to more resilient AI systems capable of continuous learning in real-world applications.
By Turing Wire Research Desk · Oct 1, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
