Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris
- Published
- Oct 1, 2026 — 17:34 UTC
{'Problem': 'The application of on-policy self-distillation to multimodal large language models (MLLMs) is largely unexplored. This paper addresses this gap by proposing a method that leverages spatially grounded guidance to enhance the performance of MLLMs.', 'Method': "The authors propose an on-policy self-distillation framework where a teacher model is provided with textual and spatially grounded guidance derived from procedurally generated scenes. These scenes include object identities and spatial coordinates, allowing the teacher to integrate evidence from multiple relevant image regions. The student model is trained to reproduce the behavior of the teacher based on the provided image and associated questions, effectively learning from the teacher's outputs in a spatially informed context.", 'Results': 'The proposed method achieves an average performance gain of 3.23 points compared to baseline models across several benchmarks, including CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld.', 'Limitations': "The authors do not report any limitations in the study. However, the lack of reported limitations may suggest a need for further exploration of the method's applicability across diverse datasets and real-world scenarios.", 'Why it matters': 'This work has significant implications for the development of more effective MLLMs by demonstrating the potential of on-policy self-distillation combined with spatially grounded guidance. It opens avenues for future research to explore similar techniques in other multimodal contexts, potentially enhancing the interpretability and performance of AI systems in complex environments.'}
By Turing Wire Research Desk · Oct 1, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
