cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents
Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck, Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh
- Published
- Sep 30, 2026 — 17:48 UTC
{'Problem': 'The paper identifies a significant gap in the reliable evaluation of the speed of computer-use agents (CUAs) due to reproducibility issues in existing benchmarks. The authors argue that current methodologies lack standardization, leading to inconsistent results across different studies. This work is presented as a preprint and has not undergone peer review.', 'Method': 'The authors propose a standardized virtual machine setup to ensure consistent testing environments. They develop a uniform execution pipeline for task evaluation, which allows for a common agent interface that facilitates seamless operation across various benchmarks. The evaluation focuses on the speed and efficiency of CUAs, analyzing four different benchmarks. Key factors examined include reasoning effort, agent harnesses, and environment latency, which are critical for understanding the performance dynamics of CUAs.', 'Results': 'The findings indicate that there is no single model family that excels in performance, speed, and cost across all evaluated benchmarks. Notably, open-weight models do not appear to be on the performance frontier. The study reveals that increasing reasoning effort can lead to faster task completion for certain models, while faster input-output operations in the environment can paradoxically slow down overall task completion time. Additionally, the authors suggest that it is possible to reduce the evaluation task set without compromising statistical power, indicating a potential for more efficient benchmarking.', 'Limitations': 'The authors do not explicitly state any limitations; however, potential issues include variability in model performance across different benchmarks, which could affect the generalizability of the results. The lack of peer review may also imply that the findings should be interpreted with caution until validated by the community.', 'Why it matters': "This work has significant implications for the field of AI benchmarking, as it provides a structured approach to evaluate CUAs' speed and efficiency. By addressing reproducibility and standardization, it lays the groundwork for future research to build upon, potentially leading to more reliable comparisons and advancements in CUA development."}
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
