Notableagents robotics

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Yifan Zhang, Yutong Dai, Viraj Prabhu, Zhiyuan Hu, Ran Xu, Zeyuan Chen

Published
Oct 5, 2026 — 17:57 UTC

Problem

The paper addresses the challenge of weak supervision in training web agents, which arises from sparse binary task success signals and the high cost of using frontier-language-model judges for evaluation. This gap in capability limits the effectiveness of web agents in real-world applications, necessitating a more robust training and evaluation framework.

Method

The authors propose CLIFT (Conformal Self-Verification), which incorporates a dual mechanism for both training and test-time operations:

  • Architecture: CLIFT employs a Compositional Conformal Certifier that retains question signals with URL-conditional evidence, matching a training-time judge. This architecture allows the agent to answer natural-language verification questions about its rollouts, enhancing the training process.
  • Training Mechanism: The agent receives signed trust weights through a polarity-aware lift, which blends the verifier score into per-step rewards without diminishing the baseline established by the judge. This approach enables the agent to learn from its own verification process, improving its decision-making capabilities.
  • Test-Time Mechanism: At test time, a certified bank is frozen and reused for Conformal Trajectory Selection (CTS). The agent samples a greedy rollout alongside diverse retries, with a self-verifier summarizing each URL trace. A conservative majority-vote rule is employed to determine whether to switch from the current incumbent agent without the need for external judge calls, streamlining the evaluation process.

Results

The proposed CLIFT framework demonstrates state-of-the-art performance across several benchmarks:

  • WebArena Infinity: Achieves leading performance among open-source web agents.
  • VisualWebArena: The bank trained with open model transfers to GPT-5.5 at test time also achieves state-of-the-art performance.
  • Online Mind2Web: The certified question bank significantly enhances the performance of live-web agents in zero-shot evaluations.

Limitations

The authors do not report any limitations in the study. However, the absence of reported limitations may suggest a need for further exploration of potential weaknesses in diverse real-world scenarios or edge cases that were not covered in the experiments.

Why it matters

The implications of this work are significant for the development of more reliable and efficient web agents. By leveraging self-verification and conformal methods, CLIFT enhances the training process and allows for more effective test-time scaling. This could lead to improved performance in various applications, including automated web interactions and information retrieval, ultimately advancing the field of AI-driven web agents.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI