Notableevaluation benchmarks

UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training

Ashish Jain, Armaan Sandhu

Published
Sep 29, 2026 — 17:17 UTC

Problem

The paper addresses a gap in the evaluation of user role execution within agent benchmarks, particularly in the context of large language models (LLMs). The authors highlight the need for a systematic approach to measure how well agents adhere to user specifications during task execution. This work is particularly relevant as it provides a structured methodology for assessing user simulators, which is crucial for improving agent performance in real-world applications. The paper is a preprint and has not yet undergone peer review.

Method

The authors propose UserProxyBench, an evaluation layer built on the tau-bench family of benchmarks. The core metric introduced is the User Fidelity Score (UFS), which quantifies the adherence of agents to user instructions. The evaluation involves 375 enterprise tasks, with the fixed agent being GPT-5.5. Key findings include a User Specification Violation Rate of 24.4% in successful episodes, indicating the frequency of deviations from user instructions. Additionally, the analysis reveals that successful episodes resulted in an average reduction of 1.06 tool calls, suggesting improved efficiency in task execution. The dominant failure mode identified is the premature disclosure of information, which highlights a critical area for further refinement in agent behavior.

Results

The results indicate a mean task reward change of 15.2 points when comparing the performance of the fixed agent (GPT-5.5) against the benchmarks established by UserProxyBench. The User Specification Violation Rate stands at 24.4% in successful episodes, underscoring the challenges agents face in fully aligning with user expectations.

Limitations

The authors do not report any limitations in their study, which may suggest a need for further exploration of potential shortcomings or areas for improvement in the UserProxyBench framework.

Why it matters

The implications of this work are significant for downstream applications in agent training and evaluation. By providing a robust framework for measuring user fidelity, UserProxyBench can enhance the development of more reliable and user-aligned LLMs. This could lead to improved performance in various applications, including customer service, personal assistants, and other interactive systems where user-agent interaction is critical.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI