PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors
Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha
- Published
- Sep 30, 2026 — 16:58 UTC
Problem
This work addresses the gap in steering pretrained generative robot policies beyond their initial effective support, particularly in out-of-distribution scenarios. The authors highlight the limitations of existing methods in adapting robot behaviors to new environments or tasks that differ from the training distribution. The paper is a preprint and has not yet undergone peer review.
Method
The authors propose a novel framework called PrefPI (Preference-Guided Policy Iteration). This framework utilizes preference-conditioned generative modeling to guide the policy updates. A key component of the method is the use of classifier-free guidance (CFG), which allows for more effective steering of the generative policies. The training data consists of 150 preference-labeled trajectories, which serve as the basis for learning the desired behaviors. The policies employed include diffusion policies and a flow-matching variant known as PI0.5. The training process is iterative, enabling the model to refine its steering capabilities over time.
Results
The results demonstrate a significant improvement in the object transport height, achieving 19.8 cm compared to a baseline of 10.7 cm. This indicates that the PrefPI framework effectively enhances the performance of generative robot policies in terms of their ability to handle out-of-distribution tasks.
Limitations
The authors do not report any limitations in their work. However, it is important to note that the absence of reported limitations does not imply that the framework is without potential drawbacks or areas for improvement.
Why it matters
The implications of this research are substantial for the field of robotics and reinforcement learning. By enabling more effective steering of generative policies, PrefPI could facilitate the deployment of robots in diverse and dynamic environments, enhancing their adaptability and utility in real-world applications. This work lays the groundwork for future research into preference-guided methods and their integration into robotic systems.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
