Semifactual Credit-Augmented Policy Optimization
Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang
- Published
- Sep 30, 2026 — 17:59 UTC
Problem
The paper addresses the sensitivity of large language models (LLMs) to task-irrelevant prompt features, which can lead to suboptimal performance. This issue is particularly relevant in the context of policy optimization methods, where the reliance on stable token responses is critical. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose Semifactual Credit-Augmented Policy Optimization (SCAPO), a causally inspired variant of Group Relative Policy Optimization (GRPO). SCAPO incorporates a mechanism for semifactual stability into token-level credit assignment, which aims to improve the robustness of LLMs against irrelevant features in prompts. The method involves measuring token probability drift through fixed responses under semifactual interventions. During training, SCAPO reduces advantages for unstable tokens in the early stages, ensuring that stability alone does not yield additional credit. The models used for training include Qwen3-4B-Base and Qwen3-1.7B-Base, although the specific training compute resources are not disclosed.
Results
The results indicate a significant improvement in accuracy on the AIME 2024-2026 benchmark, with SCAPO achieving a 5.63 percentage point increase over GRPO when using the Qwen3-4B-Base model, and a 4.17 percentage point increase with the Qwen3-1.7B-Base model. Additionally, SCAPO demonstrates superior performance on most evaluated mathematics benchmarks and all evaluated out-of-distribution benchmarks compared to other methods, although specific performance metrics for these benchmarks are not reported.
Limitations
The authors note a limitation of the GRPO approach, which assigns the same outcome-derived advantage to every response token. This can potentially reinforce spurious dependencies, undermining the model's ability to generalize effectively. The paper does not discuss other limitations or potential drawbacks of the SCAPO method.
Why it matters
The implications of this work are significant for downstream applications of LLMs, particularly in scenarios where robustness to irrelevant features is crucial. By enhancing token-level credit assignment, SCAPO could lead to more reliable and interpretable models, paving the way for improved performance in complex tasks that require nuanced understanding and reasoning.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
