EviRover: Reinforcing Agentic Perception Beyond a Glance
Kaixuan Fan, Kaituo Feng, Tianshuo Peng, Yilei Jiang, Manyuan Zhang, Junke Wang, Xiangyu Yue
- Published
- Sep 30, 2026 — 17:30 UTC
Problem
The paper addresses the challenge of perception under insufficient evidence, a critical gap in the literature that affects the reliability of AI systems in real-world applications. This work is presented as a preprint and has not undergone peer review.
Method
The authors introduce EviRover, a model architecture comprising 4 billion parameters, designed to improve perception capabilities through a two-phase training approach. The first phase involves supervised fine-tuning on a dataset generated via the EviRover-SFT-5K pipeline, followed by a reinforcement learning phase using the EviRover-RL-12K dataset. The model is evaluated on the EviLens benchmark, which consists of 688 instances across five distinct perception categories, allowing for a comprehensive assessment of its performance.
Results
EviRover demonstrates a substantial performance enhancement on the EviLens benchmark, achieving an average improvement of 30 points compared to the backbone model. Additionally, it shows a 15-point improvement on the BrowseComp-VL benchmark, indicating its effectiveness in various perception tasks.
Limitations
The authors do not report any limitations in their study, and no obvious limitations are identified in the available text.
Why it matters
The advancements presented in EviRover have significant implications for downstream applications that require robust perception capabilities in environments with limited evidence. By reinforcing agentic perception, this work paves the way for more reliable AI systems that can operate effectively in uncertain conditions.
By Turing Wire Research Desk · Sep 30, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
