UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
Ashish Jain, Armaan Sandhu
- Published
- Sep 29, 2026 — 17:17 UTC
Problem
The paper addresses a gap in the evaluation of user role execution within agent benchmarks, particularly in the context of large language models (LLMs). The authors highlight the need for a systematic approach to measure how well agents adhere to user specifications during task execution. This work is particularly relevant as it provides a structured methodology for assessing user simulators, which is crucial for improving agent performance in real-world applications. The paper is a preprint and has not yet undergone peer review.
Method
The authors propose UserProxyBench, an evaluation layer built on the tau-bench family of benchmarks. The core metric introduced is the User Fidelity Score (UFS), which quantifies the adherence of agents to user instructions. The evaluation involves 375 enterprise tasks, with the fixed agent being GPT-5.5. Key findings include a User Specification Violation Rate of 24.4% in successful episodes, indicating the frequency of deviations from user instructions. Additionally, the analysis reveals that successful episodes resulted in an average reduction of 1.06 tool calls, suggesting improved efficiency in task execution. The dominant failure mode identified is the premature disclosure of information, which highlights a critical area for further refinement in agent behavior.
Results
The results indicate a mean task reward change of 15.2 points when comparing the performance of the fixed agent (GPT-5.5) against the benchmarks established by UserProxyBench. The User Specification Violation Rate stands at 24.4% in successful episodes, underscoring the challenges agents face in fully aligning with user expectations.
Limitations
The authors do not report any limitations in their study, which may suggest a need for further exploration of potential shortcomings or areas for improvement in the UserProxyBench framework.
Why it matters
The implications of this work are significant for downstream applications in agent training and evaluation. By providing a robust framework for measuring user fidelity, UserProxyBench can enhance the development of more reliable and user-aligned LLMs. This could lead to improved performance in various applications, including customer service, personal assistants, and other interactive systems where user-agent interaction is critical.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
