Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein, Martin Gubri, Seong Joon Oh
- Published
- Oct 6, 2026 — 17:57 UTC
Problem
The paper addresses a significant gap in the capability of large language models (LLMs) regarding the high cost associated with querying them for extensive workloads. This issue is particularly relevant in scenarios where budget constraints limit the feasibility of deploying LLMs at scale. The authors propose a novel approach to mitigate these costs by evaluating the potential of LLM agents to convert their capabilities into more affordable, scalable artifacts. Notably, this work is presented as a preprint and has not undergone peer review.
Method
The authors introduce the BOTTLED benchmark, designed to assess the completion of unlabelled workloads under fixed constraints of time, compute resources, and LLM API budgets. They evaluate ten different models across three distinct tasks, focusing on zero-shot task performance and the ability to 'bottle' capabilities effectively. The evaluation metric employed is the zero-shot task performance alongside the bottling capabilities of the models. Two small-model distillation baselines are used for comparison, providing a reference point for assessing the performance of the tested models.
Results
The results indicate that 48 out of 60 bottling runs fall below the lower bound of the 95% confidence interval for zero-shot performance, suggesting a significant underperformance in many instances. Specifically, 31 of the 60 runs do not meet or exceed the performance of the stronger of the two small-model distillation baselines. In terms of query-product relevance classification, the Opus 5 model demonstrates notable efficiency, retaining approximately 82% of zero-shot macro-F1 performance while operating at a cost that is 657 times lower than traditional querying methods. Furthermore, when compared to the Jev model, Opus 5 recovers about 94% of Jev's macro-F1 performance at only a quarter of the projected full-workload cost.
Limitations
The authors acknowledge that strong zero-shot task performance does not consistently correlate with robust bottling capabilities, indicating a potential disconnect between these two aspects. Additionally, there is variability in performance among models that exhibit similar zero-shot scores, which could complicate the selection of models for specific tasks or workloads.
Why it matters
This research has significant implications for the deployment of LLMs in cost-sensitive environments, particularly in applications requiring large-scale processing of unlabelled data. By demonstrating the feasibility of bottling LLM capabilities, the work opens avenues for more efficient use of LLMs, potentially leading to broader adoption in various industries. The findings encourage further exploration into optimizing LLM performance while minimizing operational costs, which is crucial for scaling AI solutions.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
