Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan
- Published
- Sep 30, 2026 — 17:26 UTC
Problem
The paper addresses a gap in the capability of World Action Models (WAMs) to effectively reuse action experiences across different manipulation tasks. Existing WAMs face challenges in capturing cross-task semantic relationships due to background interference, which limits their applicability in diverse scenarios. This work is presented as a preprint and has not undergone peer review.
Method
The authors propose the Action Experience Dictionary (AED), which encodes historical physical action trajectories into shared action embeddings. The AED serves as a repository for action experiences, allowing for the retrieval of relevant action embeddings. An Action Tokenizer, which is a pretrained model, is utilized to extract these embeddings from the AED. To enhance the model's performance, a Cross-Attention Mechanism is implemented, which visually conditions pooled embeddings and prepends them to noisy action tokens. Additionally, a Motion-Aware Transition Loss is introduced to supervise the prediction of visual feature changes over random temporal intervals, thereby improving the model's ability to generalize across tasks.
Results
The effectiveness of the Action Experience Dictionary is verified through simulation benchmarks and real-world cross-embodiment settings. However, the available text does not report quantitative results or specific baselines against which the AED's performance is measured.
Limitations
The authors do not explicitly state any limitations in their work. However, the lack of reported quantitative results and comparisons to established baselines may hinder the assessment of the AED's performance relative to existing methods.
Why it matters
This research has significant implications for the development of more robust and flexible WAMs that can leverage historical action data to improve performance across various manipulation tasks. By addressing the reuse of action experiences, this work could pave the way for advancements in robotic manipulation and other applications where cross-task learning is essential.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
