OSWorld-Pro: Process-based Evaluation for Computer Use Agents
Zhilin Wang, Shaokun Zhang, Yifan Zhang, Hao Zhang, Jin Xu, Binfeng Xu, Jian Hu, Yunheng Zou, Karan Sapra, Andrew Tao, Jan Kautz, Yi Dong
- Published
- Sep 21, 2026 — 16:55 UTC
{'Problem': 'Current evaluation methods for Computer-Use Agents (CUAs) lack transparency regarding the reasons for task failures. This paper addresses this gap by proposing OSWorld-Pro, a process-based evaluation framework that provides insights into the specific subgoals that CUAs fail to achieve. The work is presented as a preprint and has not undergone peer review.', 'Method': 'The authors developed a comprehensive task set comprising over 300 tasks, which include more than 2800 subgoals. The evaluation framework is grounded in over 67,000 human annotations, ensuring a robust basis for assessing agent performance. The evaluation method employs human-aligned LLM-Judges to evaluate the fulfillment of subgoals, allowing for a nuanced understanding of agent capabilities and shortcomings.', 'Results': 'The performance of the evaluated agent, Claude Opus 5, was reported at 75.7% on the OSWorld-Pro benchmark, compared to a higher performance of 83.4% on the original OSWorld benchmark. This indicates a notable difference in performance metrics between the two evaluation frameworks.', 'Limitations': 'The authors identify specific failure modes that CUAs encounter, including subgoal-irrelevant actions and click-based mistakes. These limitations highlight areas where CUAs may struggle, but the paper does not discuss additional limitations or potential biases in the evaluation process.', 'Why it matters': 'The introduction of OSWorld-Pro has significant implications for the evaluation of CUAs, as it provides a clearer framework for understanding task failures. This transparency can guide future research in improving agent design and training methodologies, ultimately leading to more effective and reliable CUAs.'}
By Callan Zhang · Sep 21, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
