Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S
Christopher J. Chanhnourack
- Published
- Sep 29, 2026 — 17:04 UTC
Problem
This work addresses the evaluation of an auditable long-term memory system, focusing on the effectiveness of retrieval mechanisms in maintaining and accessing knowledge over extended periods. The paper is a preprint and has not undergone peer review, which may affect the reliability of its findings.
Method
The authors propose a hybrid candidate retrieval system that integrates several components: a cross-encoder for reranking candidates, a coverage-first packet compilation strategy, and deterministic reasoning scaffolds. The system utilizes a large language model (LLM) as a replaceable final reader, specifically employing Claude Opus accessed through an unpinned CLI alias. The candidate pool includes all gold sessions for 468 out of 470 answerable questions, and gold-complete packets were produced for 462 out of 470 questions. The evaluation process involved two passes over a set of 500 questions, scoring 479 out of 500 and 475 out of 500 under the GPT-4o model. Additionally, 72 answerable rows were updated with knowledge, using a modified scoring prompt. The comparison reader, Grok-4.6-high, achieved scores of 476 out of 474, while a maximum-reasoning-effort variant regressed to 461 out of 465. The judging agreement between two evaluators was high, with a 98.6% agreement rate on 493 out of 500 rows. Negative controls were also implemented, revealing that a verifier which repaired three incorrect drafts inadvertently broke eleven correct drafts. All components were developed using the same 500 questions, with no held-out evaluation or independent human adjudication.
Results
The system achieved the following scores:
- 479/500 and 475/500 on two evaluation passes using GPT-4o.
- 476/474 when compared to the Grok-4.6-high reader.
- 461/465 against the maximum-reasoning-effort agentic variant.
- An agreement rate of 98.6% with a second judge on 493/500 rows.
- An official judge scored both passes at 472/500. The available text does not report quantitative results beyond these scores.
Limitations
The authors note several limitations, including the absence of held-out evaluation and independent human adjudication, which raises concerns about the generalizability of the results. The sources for the retrieval and scaffold methods, as well as transcript-derived audits, are not disclosed. The headline reader received additional operator context, and complete requests were not retained, which may affect the reproducibility of the results. Furthermore, the availability of the MCP tool remains unresolved.
Why it matters
This research contributes to the understanding of auditable long-term memory systems, particularly in the context of retrieval mechanisms and their effectiveness in knowledge management. The high scores achieved on the LongMemEval-S benchmarks suggest potential for practical applications in AI systems requiring robust memory capabilities. Future work could build on these findings to enhance the reliability and transparency of memory systems in various AI applications.
By Turing Wire Research Desk · Sep 29, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
