Notableevaluation benchmarks

BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals

Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen, Esteban Gaillac, Patrick Ruch

Published
Sep 29, 2026 — 16:52 UTC

Problem

This work addresses the gap in predicting confidence from agentic RAG (Retrieval-Augmented Generation) pipeline signals, specifically within the context of the NTCIR-19 R2C2 task. The authors highlight the need for improved confidence estimation mechanisms in information retrieval systems, particularly when dealing with complex datasets such as movie corpora. The paper is a preprint and has not undergone peer review.

Method

The proposed method employs an agentic pipeline designed for searching, reading, and recording evidence from a movie corpus. The architecture integrates an entailment cascade that checks claims against cited passages to enhance the reliability of the information retrieved. The confidence calculation is orchestrated based on the remaining evidence available, although the specific model used for this calculation is not disclosed. The training compute resources utilized for this approach are also unspecified, which may limit reproducibility.

Results

The results demonstrate the performance of the proposed method against established baselines:

  • nDCG@20: 0.0709, compared to the baseline of pooled runs.
  • Accuracy: 0.9219, ranking 6th out of 25 participants.
  • HMR (Harmonic Mean Rank): 0.4915, placing 13th.
  • Revised Accuracy: 0.9375, improving the rank to 5th.
  • Revised HMR: 0.6985, achieving 9th place.
  • accHMR: 0.6549, also ranking 5th. These metrics indicate a competitive performance, particularly in accuracy and revised accuracy, suggesting that the method effectively leverages the agentic pipeline for confidence prediction.

Limitations

The authors note that the ranking on HMR can inadvertently reward low-confidence answers that are incorrect, which may skew the evaluation of the system's performance. Additionally, they propose future work that involves fitting a model on existing pipeline outputs rather than relying on manual rule crafting, indicating a potential area for improvement in the methodology.

Why it matters

This research has significant implications for downstream work in information retrieval and confidence estimation, particularly in complex domains where the reliability of retrieved information is critical. By enhancing the understanding of how to predict confidence from RAG pipeline signals, this work could lead to more robust systems capable of providing higher-quality information retrieval in various applications.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI