Notablemultimodal

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Yijia Fan, Ziqi Huang, Zhongang Cai, Yan Li, Zimo Wen, Wanqi Yin, Haiwen Diao, Ziwei Liu

Published
Sep 28, 2026 — 17:59 UTC

Problem

Unified multimodal models require the capability to jointly learn reflection text and image generation to facilitate effective self-repair. This paper addresses this gap by proposing a novel approach that integrates reinforcement learning into the training of these models. The work is presented as a preprint and has not undergone peer review.

Method

The authors introduce the UMM-Reflection model, which employs reinforcement learning (RL) to optimize the completion of reflection trajectories. The model utilizes the BAGEL dataset as its data source. The training mechanism involves sibling trajectories that share a common initial image, allowing for a comparative analysis of reflection strategies through group-relative advantage. This approach enables trajectory-level advantage updates, which refine reflection tokens and facilitate flow-based revisions. The architecture is designed to enhance the model's ability to generate coherent and contextually relevant reflections across modalities.

Results

The proposed method demonstrates significant improvements over the standard fine-tuning (SFT) baseline across multiple benchmarks:

  • GenEval: 12.05 points improvement vs SFT
  • WISE: 10.97 points improvement vs SFT
  • OneIG-Bench: 3.48 points improvement vs SFT
  • T2I-CompBench++: 4.63 points improvement vs SFT
    These results indicate the effectiveness of the UMM-Reflection model in enhancing multimodal reflection generation capabilities.

Limitations

The authors do not report any limitations in their work. However, as a preprint, the lack of peer review may imply potential oversights or unaddressed issues that could be identified in a formal review process.

Why it matters

The implications of this research are significant for downstream applications in multimodal AI systems, particularly in enhancing the self-repair capabilities of models that generate and reflect on both text and images. By improving the integration of reflection mechanisms, this work paves the way for more robust and contextually aware AI systems, which can lead to advancements in areas such as interactive AI, content generation, and automated reasoning.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI