Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
Kunlun Zhu, Xuyan Ye, Yibo Li, Cheng Qian, Beibin Li, Heng Ji
- Published
- Sep 30, 2026 — 16:40 UTC
Problem
This work addresses the gap in leveraging unsuccessful LLM agent rollouts for learning, specifically focusing on the lack of structured datasets that capture error-diagnosis pairs. The authors present the Agent Error Dataset (AED), which is designed to facilitate failure analysis and enhance error-aware post-training techniques. This preprint is particularly relevant as it provides a novel resource for researchers and engineers working on improving the robustness of language models.
Method
The core contribution is the Agent Error Dataset, which comprises 50,228 error-diagnosis pairs derived from 9,961 source tasks across 33 environments, 19 harness families, and 23 policy models. The authors propose a five-stage Agentic Error-to-Training (AET) pipeline that systematically collects failures, generates corresponding diagnoses, and validates these against recorded evidence. The methodology includes a replay comparison where corrections are evaluated against original-action retries from the same checkpoint under matched execution settings. Additionally, the training process incorporates separate views for diagnosis and actor recovery, allowing for targeted fine-tuning on 1,656 source tasks, which enhances the model's ability to recover from errors effectively.
Results
The results demonstrate significant improvements in performance metrics post fine-tuning. The verifier pass rate achieved is 51.1%, a notable increase from the original pass rate of 18.4%. The exact-step agreement improved to 63.6%, up from 47.2% before fine-tuning. The strongest prompted reference yielded a pass rate of 54.7%. Furthermore, the mean agreement improvement correlates positively with the training-set sizes. In a comparative analysis of actor training methods, action-only repair training outperformed success-only training by 6.67 percentage points on the WebShop-lite benchmark.
Limitations
The authors do not report any limitations in their study, which may suggest a need for further exploration of potential shortcomings or biases in the dataset or methodology. However, the absence of reported limitations could also indicate a focus on the dataset's strengths and immediate applicability.
Why it matters
The introduction of the Agent Error Dataset and the AET pipeline has significant implications for downstream work in the field of AI and machine learning. By providing a structured approach to analyzing and learning from agent failures, this research paves the way for more robust error-aware training methodologies. It encourages the development of models that can better handle failures, ultimately leading to more reliable and effective AI systems.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
