動画がプログラム指定の出来事を描いているかを評価
ROWBench: Do Video Models Render What the Program Specifies?
この論文をやさしく読む
ひとことで言うと
生成動画がきれいに見えるかだけでなく、プログラムで指定した出来事を正しい順序で描いているかを測るベンチマークです。画面外で起きたことも記録し、後で見える結果と照合します。
何に役立つ?
想定用途は、ゲーム向けなどの世界モデルで、物体の動作や相互作用、長い時間の整合性を評価することです。プログラムの記録を映像評価の根拠にできます。
この研究の面白いところ
同じ出来事から複数の視点や粗い3D、枠による表現を生成できるため、制御情報と出力映像を対応づけられます。視野外の出来事も記録対象にしています。
どこまで分かった?
要旨はベンチマークの構成と評価方法を説明しており、モデル別の成績は報告していません。複数視点があるのはエピソードの一部です。タイトルではROWBench、要旨ではPROWBenchと表記されており、全文訳では要旨の表記を保持しています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
プログラム可能な世界モデルは、実行可能な動力学と視覚生成を分離し、次世代ゲームエンジンの有望な基盤となる。しかし、明示された規則や相互作用に映像が従っているかは、十分に評価されていない。既存のベンチマークは視覚品質、制御可能性、指示や物理への準拠を評価するが、プログラムが指定した細かな世界内の出来事に対する忠実さを検査することは少ない。 本研究では、さまざまな場面と相互作用を含む、プログラムで構築した170エピソードと600本の代理動画からなるPROWBenchを導入する。PROWBenchは、カメラの視野外のものも含め、実体の状態と時刻付きの出来事を再生可能な世界記録として保存し、そこから同期した視点映像と代理表現を描画する。これにより、プログラム実行がもたらす観測可能な結果に照らして、生成動画を検査できる。拡張可能な枠組みにより場面を構築し、振る舞いを制御し、各カメラ視点を粗い3Dやバウンディングボックスなどの異なる表現で描画できる。 ベンチマークは一人称と三人称の視点を含み、一部のエピソードでは同期した複数視点の観測も利用できる。PROWBenchはこれらの記録を根拠として、実体の制御と長期的な記憶を評価する。さらに、VLMに基づく二つの指標、Logic-Render AlignmentとInteraction Success Rateを用いて、指定された時系列への準拠と、エンジンが時刻付きで記録した出来事が視覚的に実現されているかを評価する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.
arXiv ID: 2610.02205 / 要約の誤りについて