Notablemultimodal

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang, Shiyu Chang

Published
Sep 30, 2026 — 17:55 UTC

Problem

The paper addresses a gap in the capability of existing systems to efficiently identify anomalies in interactive 3D worlds. It highlights the need for robust pipelines that can handle complex environments and tasks, particularly in the context of multimodal agents. The work is presented as a preprint, indicating that it has not yet undergone peer review.

Method

The authors propose the WorldAuditBench benchmark, which consists of 213 anomaly detection tasks across 13 environments created using Unreal Engine 5 and Three.js. The benchmark evaluates five different models against five distinct families of anomalies. The auditing paradigms employed include:

  • A Visual Language Agent (VLA)-based exploration followed by a Visual Language Model (VLM)-based anomaly identification.
  • An end-to-end VLM agent that utilizes visual reasoning to guide action selection. The success rates of the models evaluated range from 6.6% to 42.3%, with a human performance benchmark set at 83.4%. This disparity underscores the challenges faced by current multimodal agents in achieving human-level performance in anomaly detection tasks.

Results

The success rates of the evaluated models fall between 6.6% and 42.3%, significantly lower than the human performance benchmark of 83.4%. This indicates that while the models can perform anomaly detection, they are still far from achieving the efficiency and accuracy of human auditors in interactive 3D environments.

Limitations

The authors note that current multimodal agents exhibit limitations in their ability to gather and interpret evidence during exploration. This suggests that further advancements are needed in the design of these agents to enhance their performance in complex environments. Additionally, the paper does not address potential scalability issues or the generalizability of the models across different types of 3D environments beyond those tested.

Why it matters

The introduction of WorldAuditBench has significant implications for future research in the field of anomaly detection and multimodal learning. By providing a standardized benchmark, it facilitates the comparison of different models and approaches, potentially accelerating advancements in the development of more capable agents. The findings highlight the current limitations of multimodal agents, which can inform future work aimed at bridging the performance gap between automated systems and human auditors.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI