A Living Benchmark for Information Retrieval from Electronic Health Records
Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani, Philip Chung, Kevin R Keet, Kameron C. Black, Andrea T. Fisher, Sarita Khemani, Jerry Liu, Stephen Ma, Saloni K. Maharaj, Rita M. Pandya, Eduardo Perez-Guerrero, Priyanka Pillai, Lisa Shieh, David J. H. Wu, James Xie, James C. McAvoy, Teresa Nguyen, Jessica Tran, Lucy Yin, Bridget Lin, Alison Callahan, Jason A. Fries, Nigam H. Shah, Emily Alsentzer
- Published
- Sep 24, 2026 — 17:41 UTC
{'Problem': 'Existing benchmarks for evaluating large language models (LLMs) in electronic health records (EHRs) are manually curated, expensive to maintain, and rapidly become outdated. This paper addresses the need for a more adaptive and relevant evaluation framework that can keep pace with the evolving nature of clinical data and the requirements of information retrieval tasks in healthcare.', 'Method': 'The authors propose a framework that automatically generates question-answer pairs from longitudinal EHR notes, creating a new dataset called the Benchmark for Retrieving Information in EHRs (BRIE). The benchmark generator is validated by nineteen clinicians, ensuring that the generated content is clinically relevant and accurate. This approach allows for the continuous updating of the benchmark, making it more sustainable and reflective of current clinical practices.', 'Results': 'The BRIE benchmark reveals that state-of-the-art systems often omit clinically important information when evaluated against static benchmarks. Additionally, BRIE enables evaluations that static benchmarks cannot support, such as the generation of multiple answers that reflect clinician reasoning, thereby providing a more nuanced assessment of LLM performance in EHR contexts.', 'Limitations': "The authors do not report any limitations in the study. However, the reliance on clinician validation may introduce subjectivity, and the framework's effectiveness in diverse clinical settings remains to be tested.", 'Why it matters': 'This work has significant implications for downstream research in medical informatics and AI in healthcare, as it provides a living benchmark that can adapt to new clinical insights and practices. It encourages the development of more robust LLMs capable of understanding and retrieving critical information from EHRs, ultimately improving patient care and clinical decision-making.'}
By Callan Zhang · Sep 24, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
