DepthWorld: 3D World Model for Robot Manipulation
Jai Bardhan, Josef Sivic, Vladimir Petrik
- Published
- Oct 6, 2026 — 17:59 UTC
Problem
This work addresses a gap in the fidelity of 3D geometry in video-based world models for robotics. The authors highlight the limitations of existing models in accurately capturing depth information, which is crucial for effective robot manipulation tasks. The paper is a preprint and has not undergone peer review.
Method
The authors introduce a novel calibration pipeline that integrates learned stereo depth with a joint factor graph to enhance depth estimation. The data source utilized is the DROID dataset, which is leveraged to produce DROID-3D, a calibrated 3D dataset featuring dense metric depth and recalibrated multi-view extrinsics. The model, named DepthWorld, is based on a Stable Video Diffusion architecture. It employs a prediction mechanism that jointly predicts multi-view RGB and depth through spatial latent tiling. Notably, a Variational Autoencoder (VAE) component is incorporated into the model but remains unchanged from its original configuration. The reprojection error achieved is less than 0.7 pixels on 90% of episodes for external cameras, indicating high accuracy in depth estimation.
Results
The proposed DepthWorld model demonstrates a Peak Signal-to-Noise Ratio (PSNR) improvement of +1.48 dB over an RGB-only baseline while maintaining an equal training budget. This improvement signifies enhanced performance in depth perception compared to traditional RGB-only approaches.
Limitations
The authors do not report any limitations in their work. However, as with many models relying on specific datasets, the generalizability of the results to other environments or datasets may warrant further investigation.
Why it matters
The implications of this research are significant for downstream applications in robotic manipulation, where accurate depth perception is critical. By improving the fidelity of 3D geometry in world models, DepthWorld could enhance the performance of robots in complex environments, potentially leading to advancements in autonomous navigation and interaction tasks.
By Turing Wire Research Desk · Oct 6, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
