FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
- Published
- Oct 1, 2026 — 17:59 UTC
Problem
Critics trained to predict returns may not yield accurate action derivatives for policy updates. This issue is particularly relevant in reinforcement learning contexts where the quality of the critic can significantly impact the performance of the policy. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose Forward Entropy-Regularized Policy Optimization (FERPO), which aims to improve policy performance by utilizing critic values without the need to differentiate the critic with respect to actions. Key components of the method include:
- Regularization: The algorithm employs entropy and Kullback-Leibler (KL) divergence to regularize the policy updates.
- Target Action Distribution: This is derived from a policy-improvement objective, ensuring that the updates are aligned with the desired policy behavior.
- Loss Function: The forward-KL objective is utilized to guide the optimization process.
- Estimation Method: Self-normalized importance sampling (SNIS) is employed, with actions drawn from the rollout policy, allowing for effective estimation of the expected returns.
- Training Environment: The experiments are conducted in the MuJoCo Playground and ManiSkill environments, which are standard benchmarks for evaluating reinforcement learning algorithms.
Results
The results indicate that FERPO achieves competitive performance and sample-efficiency gains, although specific quantitative results are not reported. Additionally, the actor update speed of FERPO is noted to be faster than that of Relative Entropy Pathwise Policy Optimization (REPPO), suggesting improvements in the efficiency of policy updates compared to this established baseline.
Limitations
The authors do not report any limitations in their work, and no obvious limitations are identified in the provided text.
Why it matters
The implications of this work are significant for downstream applications in reinforcement learning, particularly in scenarios where sample efficiency is critical. By addressing the limitations of traditional critic-based methods and providing a novel approach to policy optimization, FERPO may facilitate more effective learning in complex environments, potentially leading to advancements in various applications of reinforcement learning.
By Turing Wire Research Desk · Oct 1, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
