GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
Yiran Wang, Xingyilang Yin, Junfu Pu, Guangzhi Wang, Kaifeng Li, Mingyu Ouyang, Huiqiang Sun, Lingen Li, Cheng Cheng, Wangbo Yu, Honghao Chen, Xiaodong Cun, Chi-Man Pun, Zhiguo Cao, Ying Shan
- Published
- Sep 21, 2026 — 17:59 UTC
Problem
The paper addresses a significant gap in existing datasets and benchmarks for AI models in video games, particularly in the context of multi-horizon gameplay evaluation. The authors highlight the need for a structured approach to assess AI performance across various gameplay scenarios, which is currently lacking in the literature. This work is presented as a preprint, indicating that it has not yet undergone peer review.
Method
The authors propose the GameHorizon Suite, which consists of three main components:
- GameHorizon-Annotator: A scalable and automated annotation pipeline designed to generate multi-horizon instructions for gameplay data.
- GameHorizon-Data: A large-scale dataset comprising 5,000 hours of gameplay recordings from 21 AAA games. This dataset was collected by 100 human expert players and includes temporally aligned videos, player actions, and multi-horizon instructions, providing a rich resource for training and evaluating AI models.
- GameHorizon-Bench: A reproducible testing framework that includes both offline and stepwise online evaluation tracks. The offline track utilizes thousands of standardized questions organized into three primary tasks and diagnostic variants to assess model performance. The online track correlates offline scores with actual gameplay capabilities, allowing for the localization of failures in long-horizon gameplay scenarios.
Results
The authors evaluated 47 models through over one million model invocations, revealing a hierarchy of task difficulty and significant differences in model capabilities. However, the available text does not report quantitative results such as specific performance metrics or comparisons against baseline models.
Limitations
The authors do not report any limitations in their work. However, the absence of reported limitations may suggest a lack of critical self-assessment or acknowledgment of potential challenges in the dataset or evaluation framework.
Why it matters
The GameHorizon Suite has significant implications for downstream work in AI for video games. By providing a comprehensive dataset and evaluation framework, it enables researchers to develop and benchmark AI models more effectively across diverse gameplay scenarios. This could lead to advancements in AI capabilities in gaming, enhancing both the development of intelligent agents and the overall gaming experience.
By Callan Zhang · Sep 21, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
