Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage
Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
- Published
- Sep 24, 2026 — 17:45 UTC
Problem
This work addresses the gap in predictive modeling for underwater Remotely Operated Vehicle (ROV) salvage operations, specifically in scenarios where contact sensors are not utilized. The authors highlight the challenges in accurately predicting object states in underwater environments, which are often characterized by complex visual conditions and the need for robust task execution.
Method
The proposed architecture, Underwater C$^{3}$-JEPA, is a cross-view, control-conditioned, context-extended model designed to predict task-object state evolution in latent space. Key components of the method include:
- Encoding: The model encodes synchronized multi-view RGB observations into task-object and context tokens, allowing for a rich representation of the environment.
- Attention Mechanism: It employs held-out-view attention to fuse cross-camera evidence, enhancing the model's ability to leverage information from multiple perspectives.
- State Prediction: Future states are predicted based on vehicle control signals, enabling the model to adapt to dynamic underwater conditions.
- Weak Binding Anchors: The architecture incorporates weak binding anchors for the target and gripper, facilitating better alignment of the model's predictions with the physical task.
- SIGReg Component: A sharp geometric representation is achieved through the SIGReg component, which refines the model's understanding of spatial relationships.
Training compute details are not specified in the paper.
Results
The results indicate that the learned representation from Underwater C$^{3}$-JEPA transfers more task-relevant information to downstream probes compared to a reconstruction-free latent baseline. Additionally, the architecture demonstrates superior performance in recovering a withheld camera's object state when validated against persistence in real underwater video scenarios. The specific baselines compared are the reconstruction-free latent baseline and persistence, respectively.
Limitations
The authors do not report any limitations in their work. However, the absence of training compute specifications may limit reproducibility and scalability assessments.
Why it matters
The implications of this research are significant for advancing underwater robotics, particularly in enhancing the autonomy and efficiency of ROVs in salvage operations. By improving predictive modeling without reliance on contact sensors, this work opens avenues for more robust underwater exploration and task execution, potentially leading to broader applications in marine research and recovery missions.
By Callan Zhang · Sep 24, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
