Notableagents robotics

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

Published
Sep 30, 2026 — 16:58 UTC

Problem

This work addresses the gap in steering pretrained generative robot policies beyond their initial effective support, particularly in out-of-distribution scenarios. The authors highlight the limitations of existing methods in adapting robot behaviors to new environments or tasks that differ from the training distribution. The paper is a preprint and has not yet undergone peer review.

Method

The authors propose a novel framework called PrefPI (Preference-Guided Policy Iteration). This framework utilizes preference-conditioned generative modeling to guide the policy updates. A key component of the method is the use of classifier-free guidance (CFG), which allows for more effective steering of the generative policies. The training data consists of 150 preference-labeled trajectories, which serve as the basis for learning the desired behaviors. The policies employed include diffusion policies and a flow-matching variant known as PI0.5. The training process is iterative, enabling the model to refine its steering capabilities over time.

Results

The results demonstrate a significant improvement in the object transport height, achieving 19.8 cm compared to a baseline of 10.7 cm. This indicates that the PrefPI framework effectively enhances the performance of generative robot policies in terms of their ability to handle out-of-distribution tasks.

Limitations

The authors do not report any limitations in their work. However, it is important to note that the absence of reported limitations does not imply that the framework is without potential drawbacks or areas for improvement.

Why it matters

The implications of this research are substantial for the field of robotics and reinforcement learning. By enabling more effective steering of generative policies, PrefPI could facilitate the deployment of robots in diverse and dynamic environments, enhancing their adaptability and utility in real-world applications. This work lays the groundwork for future research into preference-guided methods and their integration into robotic systems.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI