Notableevaluation benchmarks

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn

Published
Oct 1, 2026 — 17:59 UTC

Problem

This work addresses a gap in the capability of AI systems to effectively identify relevant prior research that can inspire solutions to new problems. The authors highlight the inadequacy of existing retrieval methods in the context of academic literature, particularly in computer science, and propose a new benchmark to evaluate these capabilities. The paper is a preprint and has not yet undergone peer review.

Method

The authors constructed a dataset comprising 207 recent computer science papers, which were annotated by 184 lead authors. The primary task is framed as a retrieval task, where the goal is to retrieve papers that are deemed relevant based on author-provided judgments. The evaluation metric employed is Recall@20, which measures the proportion of relevant papers retrieved within the top 20 results. This metric is critical for assessing the effectiveness of retrieval systems in academic contexts.

Results

The paper reports the following results:

  • The embedding retrieval method achieved a Recall@20 score of 0.48, outperforming the baseline method, Agentic Search, which scored 0.42.
  • The Claude Fable 5.1 model achieved a Recall@20 score of 0.51, although the baseline for this comparison is not specified in the text. These results indicate a promising direction for improving retrieval systems in academic research.

Limitations

The authors acknowledge the need for new training recipes to enhance the search capabilities of models. This suggests that while the current methods show some promise, there is significant room for improvement in model training and architecture to achieve better retrieval performance. Additionally, the dataset's reliance on a limited number of annotated papers may restrict the generalizability of the findings.

Why it matters

The implications of this work are substantial for downstream research in AI and information retrieval. By establishing a benchmark like ScholarCatalyst, the authors provide a framework for future studies to evaluate and improve the ability of AI systems to connect researchers with relevant literature. This could lead to more effective research methodologies and foster innovation by ensuring that researchers can easily access and build upon existing work.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI