VideoTapestry: Query-Adaptive Memory Refinement for Multi-Agent Long-Video Understanding
Yucheng Liu, Yufei Yin, Mingxiao Feng, Jiajun Deng, Wengang Zhou, Houqiang Li
- Published
- Oct 5, 2026 — 16:44 UTC
Problem
This work addresses the challenge of long-video understanding and memory retrieval across extended temporal spans. The authors propose a solution in the form of a training-free multi-agent system, which is particularly relevant given the increasing complexity of video data and the limitations of existing models in handling long sequences effectively.
Method
The proposed framework, VideoTapestry, utilizes a hierarchical video memory structure comprising three levels: 1) global narrative context, 2) event-level temporal structure, and 3) fine-grained relational evidence. Each level is managed by specialized agents that facilitate coarse-to-fine localization and observation. The refinement process involves these agents revisiting relevant video regions to enrich the memory with multimodal observations, ultimately assembling a composite query-adaptive memory. This architecture allows for dynamic adaptation to specific queries, enhancing the model's ability to understand and retrieve information from long videos without the need for extensive training.
Results
The authors report significant accuracy gains on several benchmarks compared to the baseline model GPT-5.5: 17.2% on LVBench, 14.9% on LongVideoBench (Long), 9.8% on Video-MME (Long), and 7.0% on EgoSchema. These results indicate that VideoTapestry achieves state-of-the-art performance among all competitors in the evaluated tasks, demonstrating its effectiveness in long-video understanding.
Limitations
The authors do not report any limitations in their work, and no obvious limitations are identified in the available text. This absence of reported limitations may suggest a need for further exploration of the framework's robustness across diverse video types and contexts.
Why it matters
The implications of this work are significant for downstream applications in video analysis, particularly in scenarios requiring nuanced understanding of long-form content. By improving memory retrieval and understanding in long videos, VideoTapestry could enhance various applications, including video summarization, event detection, and content-based retrieval systems, paving the way for more sophisticated AI-driven video analytics.
By Turing Wire Research Desk · Oct 5, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
