ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
Yong Du, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, Yongliang Shen
- Published
- Sep 30, 2026 — 17:40 UTC
Problem
Existing methods for computer-use agents (CUAs) primarily depend on sparse outcome rewards, which limits their ability to learn from intermediate actions. This paper addresses the lack of supervision for these intermediate actions, proposing a novel approach to enhance the learning process of CUAs. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose ComputerSD, an online self-distillation method specifically designed for CUAs. The core of the method is an on-policy self-distillation (OPSD) algorithm that integrates trajectory-level Generalized Reweighted Policy Optimization (GRPO). The training leverages real-time feedback derived from executed graphical user interface (GUI) transitions, allowing the model to learn from immediate actions rather than solely from final outcomes. The training framework is fully asynchronous, which facilitates efficient learning. Additionally, a fine-tuned GUI analyzer is employed to provide guidance and a step-level value score after each action, enhancing the agent's decision-making process.
Results
The proposed method demonstrates significant improvements over the baseline outcome-only GRPO approach. Specifically, it achieves a 1.9 percentage point improvement on the Qwen3-VL-8B-Thinking benchmark and a 4.1 percentage point improvement on the EvoCUA-8B benchmark. These results indicate the effectiveness of the ComputerSD method in enhancing the performance of CUAs through real-time feedback and self-distillation.
Limitations
The authors do not report any limitations in their work. However, the absence of reported limitations may suggest a need for further exploration of potential challenges in diverse real-world applications or scalability issues.
Why it matters
The implications of this work are significant for the development of more effective CUAs. By enabling agents to learn from real-time feedback and intermediate actions, ComputerSD could lead to more robust and adaptable systems in various applications, including automated customer service, personal assistants, and other interactive AI systems. This approach may pave the way for future research into self-distillation techniques and their application in reinforcement learning contexts.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
