Recursive Video In-Context Learning for Agentic Robot
Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang
- Published
- Oct 5, 2026 — 17:59 UTC
Problem
The paper addresses a significant gap in the capabilities of large language model (LLM) agents, specifically their ineffective integration of demonstration videos into the context of task execution. This limitation hinders the ability of agents to leverage visual information for improved performance in robotic tasks. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose a novel approach called Recursive Video In-Context Learning (RV-ICL). This method involves transforming demonstration videos into a navigable hierarchy of sub-events, which include grasps and releases. The hierarchy is structured into several levels: keyframes, phases, moments, and short clips. The data requirement for this method is minimal, necessitating only one demonstration per task. During execution, the agent first reads the coarse levels of the hierarchy to plan its actions and then re-enters the hierarchy for detailed execution, allowing for a more nuanced understanding of the task at hand.
Results
The proposed RV-ICL method demonstrates significant improvements in task success rates on benchmark datasets. On the LIBERO-PRO benchmark, the success rate achieved is 96.5%, compared to a baseline of 92.6%. Similarly, on the LIBERO-Plus benchmark, the success rate is 95.8%, surpassing the baseline of 86.7%. These results indicate a marked enhancement in the agent's ability to execute tasks effectively when utilizing the RV-ICL framework.
Limitations
The authors do not report any limitations in their work. However, it is important to note that the reliance on a single demonstration per task may limit the generalizability of the approach across diverse tasks and environments, a consideration that could be explored in future research.
Why it matters
The implications of this work are significant for the field of robotic task execution and LLM integration. By effectively incorporating visual demonstrations into the task execution context, RV-ICL has the potential to enhance the performance of robotic agents in real-world applications. This approach could pave the way for more sophisticated interaction between LLMs and visual data, leading to advancements in autonomous systems and human-robot collaboration.
By Turing Wire Research Desk · Oct 5, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
