Notablereasoning

Selective Transfer of RL Updates for Visual Reasoning

Suxin Ji, Hungtao Wan, Mingjun Liu, An Zhang

Published
Oct 6, 2026 — 16:43 UTC

Problem

The paper addresses the issue of endpoint-based transfer in reinforcement learning (RL), which conflates pre-existing model differences with updates acquired during reasoning post-training. This conflation can hinder the effective transfer of reasoning capabilities from language models to vision-language models (VLMs). The authors propose a solution to this problem, which is particularly relevant given the increasing complexity of visual reasoning tasks.

Method

The authors introduce two key components: Model Merging and Selective-RL.

  • Model Merging offers a training-free approach to transfer reasoning capabilities from language models to VLMs, allowing for a more efficient integration of learned knowledge without the need for extensive retraining.
  • Selective-RL focuses on isolating the RL-stage updates, specifically retaining dominant matrix-wise directions while preserving their magnitudes. This selective transfer mechanism enables the authors to effectively transfer these updates to the language modules of a VLM, enhancing its reasoning capabilities without the drawbacks associated with full-update interpolation.

Results

The proposed method demonstrates significant improvements in performance metrics. Specifically, on the MathVision benchmark, the Selective-RL approach achieves an improvement of 8.55 percentage points compared to full-update interpolation. Furthermore, Selective-RL outperforms full-update interpolation in 12 out of 15 comparative evaluations, indicating its robustness and effectiveness across various scenarios.

Limitations

The authors do not report any limitations in their work, suggesting that the method is robust and effective in its current form. However, the absence of reported limitations may warrant further scrutiny in practical applications, as real-world scenarios often introduce unforeseen challenges.

Why it matters

This work has significant implications for downstream applications in visual reasoning and multimodal learning. By providing a method for selective transfer of RL updates, the authors pave the way for more efficient training and deployment of VLMs, potentially leading to advancements in tasks that require complex reasoning across modalities. The approach could also inspire further research into selective transfer mechanisms in other areas of machine learning, enhancing the adaptability and performance of models in diverse applications.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI