ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
Hanjun Luo, Xiucheng Zhang, Zhuoning Xu, Zhimu Huang, Yingbin Jin, Xinfeng Li, Hanan Salam
- Published
- Oct 6, 2026 — 16:46 UTC
Problem
The paper addresses the lack of a systematic framework for the unified evaluation of risk-treatment capabilities in coding agents. This gap is particularly relevant in the context of agentic coding, where the decision-making processes of agents can lead to unnecessary defensive actions that do not align with explicit evidence. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose the Avoidance-Transfer-Mitigation-Acceptance (ATMA) framework, which is rooted in software engineering risk management principles. The evaluation is conducted using a dataset comprising 200 evidence-controlled repository-level task pairs. The authors introduce dedicated metrics to assess risk-treatment violations and evidence responsiveness, ensuring a comprehensive evaluation of agentic behavior. The evaluation method employs a human-calibrated agentic judge to provide a nuanced assessment of the agents' performance in risk treatment.
Results
The findings reveal that between 11.2% and 58.7% of runs exhibit unnecessary risk treatment when compared to explicit evidence. Furthermore, the analysis indicates that a stronger task capability does not guarantee more appropriate risk treatment, suggesting a disconnect between task performance and risk management. Additionally, the study highlights that treatment violations significantly impair the developer experience, indicating that unnecessary defensive actions can have broader implications beyond mere performance metrics.
Limitations
The authors do not report any limitations in their study. However, the absence of reported limitations may suggest a need for further exploration of the framework's applicability across diverse coding environments and agentic behaviors.
Why it matters
The implications of this work are significant for downstream research in agentic coding and risk management. By establishing a framework for evaluating risk-treatment capabilities, the study paves the way for more effective design and implementation of coding agents that can balance task performance with appropriate risk management. This could lead to improved developer experiences and more reliable coding practices in automated systems.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
