Notableevaluation benchmarks

ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences

Mingda Zhang, Wenjin Liu, Tiesunlong Shen, Zikai Xiao, Zhenghong Lin, Qing Xu, Erik Cambria, Xiaoying Tang, Haoran Luo

Published
Oct 6, 2026 — 17:08 UTC

Problem

The paper addresses a significant gap in the evaluation of persistent program-level improvements in scientific automation, particularly in the context of sequential tasks across both natural and social sciences. The authors highlight the absence of comprehensive benchmarks that assess the continual self-evolution capabilities of AI agents in scientific domains, which is critical for advancing automation in research.

Method

The authors propose a novel framework named ScienceClaw, which is designed to facilitate the evaluation of AI agents' continual self-evolution. The evaluation method, termed ScienceClaw-Eval, encompasses 23 different scientific disciplines. Key metrics measured include scientific correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost. The mechanism employed by ScienceClaw involves repairing executable workflows through multi-turn interactions, allowing the system to learn from past failures and successes. Specifically, it converts re-execution-verified failure-success trajectories into linked Skill and Operator candidates. Updates to the system are retained only when the source-task replay successfully reproduces the repair and demonstrates improvements in independent scientific tasks. The authors have made the code repository available for further exploration and implementation.

Results

The available text does not report quantitative results regarding the metrics of scientific correctness, evolutionary gain, retention, cross-dataset transfer, or evolution cost. This lack of reported results limits the ability to assess the framework's performance against existing baselines.

Limitations

The authors do not explicitly state any limitations within the paper. However, potential limitations could include the breadth of disciplines covered and the nature of the evaluations conducted, which may not fully capture the complexities of continual self-evolution in diverse scientific contexts.

Why it matters

The implications of this work are significant for downstream research in AI-for-Science. By establishing a framework for continual self-evolution, ScienceClaw could enhance the capabilities of AI agents in automating scientific processes, leading to more efficient research methodologies. Furthermore, the introduction of a structured evaluation method may pave the way for future studies to benchmark and improve AI systems in various scientific fields, ultimately contributing to advancements in scientific discovery and innovation.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI