TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
Oliver Jaffe, Dane Sherburn
- Published
- Oct 5, 2026 — 17:57 UTC
Problem
Evaluating the experimental research taste of AI systems is a critical gap in the literature, particularly in understanding how these systems compare to human experts. This work addresses this gap by proposing a novel benchmark, TasteVal, to systematically assess AI performance in experimental design and implementation tasks. The paper is a preprint and has not yet undergone peer review.
Method
The authors introduce the TasteVal benchmark, which consists of 8 novel, open-ended tasks designed to evaluate AI systems in two distinct roles: Researcher (responsible for designing experiments) and Coder (responsible for implementing experiments). The evaluation involves 20 AI models released between 2023 and 2026, with a focus on the best-performing model, Opus 5.5. The execution budgets for the tasks are set at 40 H100 hours or 120 wall-clock hours. To assess model performance, the authors recruited 24 human experts, ensuring at least 2 experts per task. The evaluation metric is a compute multiplier, which compares the AI model scores against the expert baseline.
Results
The best-performing model, Opus 5.5, achieved a compute multiplier of 2.3x (95% CI 1.15-4.37) compared to the expert baseline. Additionally, Opus 5.5 operates at approximately 1/30 of the average per-run cost of the baseline models. The compute multiplier trend indicates that AI performance has doubled every 3.0 months since December 2025 (95% CI 1.7-5.0), a significant acceleration compared to the previous doubling rate of every 14 months between 2023 and December 2025. The final normalized performance trend shows a doubling every 14.6 months.
Limitations
The authors note that the specific tasks used in the benchmark have not been released to maintain the integrity of the TasteVal evaluation, preventing potential contamination of results. This limitation may hinder reproducibility and broader applicability of the findings. Additionally, the reliance on a limited number of human experts may introduce variability in the baseline performance assessment.
Why it matters
The introduction of TasteVal has significant implications for the field of AI research, particularly in the evaluation of AI systems' capabilities in experimental design. By providing a structured framework for comparison against human experts, this benchmark can guide future research in improving AI systems' experimental research taste. The observed trends in compute multiplier performance suggest rapid advancements in AI capabilities, which could influence the development of more sophisticated AI systems in various domains.
By Turing Wire Research Desk · Oct 5, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
