NotableotherMicrosoft

ThinkingBox Evaluates 507 Workflows, Reveals 79,853 Failed Attempts

Published
Oct 3, 2026 — 22:56 UTC
Also in this story:Hugging Face

ThinkingBox Evaluates 507 Workflows, Reveals 79,853 Failed Attempts

ThinkingBox evaluated 507 stateful business workflows, running each against various large language models (LLMs) 20 times, resulting in a total of 121,680 valid trials. This evaluation revealed that 79,853 attempts failed, with 79.9% of those failures attributed to tool usage. The overall task-weighted pass@1 score for Claude Opus 5.5 was 67.16%, while Claude Opus 5 achieved a perfect 100% pass rate in the auto insurance domain.

The evaluation involved AI models including Claude Opus 5.5, GPT-5.6, and GPT-6 Astra, with GPT-5.6 achieving a 104% pass@1 score in both the travel and neobank domains. The costs associated with these evaluations were notable, with GPT-5.4 costing $43.49 for 507 attempts and $0.131 per successful task attempt. In contrast, GPT-6 Astra's estimated cost for 20 runs was $1,720.60, translating to $7.45 per dependable task.

The findings highlight a substantial gap in AI agent reliability, as evidenced by the claim that an agent processing a refund correctly once but mishandling it four times is not a functional refund agent. This follows previous reports of disturbing content in Microsoft Copilot prompts and the launch of new AI models by Microsoft, indicating ongoing challenges in AI reliability and performance evaluation. The data underscores the necessity for practitioners to prioritize dependable task performance over cost efficiency in AI deployments.

Summarised from Hugging Face Blog's original report by the Turing Wire Newsdesk. Read the original for the full story.

Source: Hugging Face Blog