ゲーム操作能力を短期から長期まで評価するデータセット
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
この論文をやさしく読む
ひとことで言うと
AIがゲーム画面を理解するだけでなく、短い操作から長い目標達成までどこまでできるかを測る共通の評価基盤です。21ゲームの人間の操作記録を使っています。
何に役立つ?
ゲーム操作モデルの比較や、長い操作手順のどこで失敗するかの診断に使うことが想定されています。オフライン得点と実際のプレイを結び付けて確認できる設計です。
この研究の面白いところ
5,000時間の動画・操作・指示を対応付け、47モデルを100万回超の呼出しで評価しています。得点を付けるだけでなく、時間の長さによって要求能力がどう変わるかを調べています。
どこまで分かった?
要旨はモデル間の違いと課題の難易度階層を報告しますが、モデル別の数値は示していません。データセットなどは公開予定とされており、この要旨だけで公開済みとは判断できません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
現代のビデオゲームは、視覚理解、指示の分解、目標計画、精密な行動制御を複数の時間的範囲にわたって組み合わせる、AIモデルの測定可能な試験環境となる。しかし既存のデータセットやベンチマークは、対象ゲームの範囲が狭い、言語による指示がない、あるいはばらつきの大きいオンライン実行に依存する、という問題を抱えている。これらの課題に対処するため、多様なモデル群について異なる時間的範囲のゲームプレイ能力を測る、統一的なデータ・評価スイートGameHorizonを導入する。 GameHorizon Suiteは3つの構成要素からなる。第1に、GameHorizon-Annotatorは、複数の時間的範囲の指示を付与する、拡張可能な自動アノテーションパイプラインである。第2に、このパイプラインを用いて、動画、プレイヤーの行動、複数の時間的範囲の指示を時間的に対応付けた、初の大規模AAAゲームプレイデータセットGameHorizon-Dataを構築する。100人の熟練した人間のプレイヤーが収集した、21ゲーム、5,000時間の記録からなる。第3に、再現可能なオフライン試験と、段階ごとのオンライン試験を備えたGameHorizon-Benchを構築する。オフライン部門は、3つの主要課題と一連の診断用派生課題に整理された数千の標準化質問を用いて、再現可能な評価を可能にする。一方オンライン部門は、オフラインの得点が実際のゲームプレイ能力を反映するかを調べ、長期のゲームプレイの特定の段階へ失敗原因を絞り込む。 GameHorizon Suiteに基づき、100万回を超えるモデル呼出しを通して47モデルを評価し、課題の難しさの意味のある階層と、モデル能力の顕著な違いを明らかにした。本研究は、時間的範囲やモデル群をまたいでゲームプレイ能力を評価する標準的な尺度を提供できる。今後の研究を促進するため、データセット、アノテータ、ベンチマークを公開する予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.
著者のコメント
We will release our dataset, annotator, and benchmark to facilitate future research. Github Repo: https://github.com/TencentARC/GameHorizon & Project Page: https://gamehorizon-suite.github.io
arXiv ID: 2609.25001 / 要約の誤りについて