arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

火災の複数の物理場からモデルの推論力を測る評価基盤

FireWorldBench: Benchmarking Complex Physical World Intelligence through Coupled-Field Fire Dynamics

Qiang Chen, Hao Guo, Huatai Zhu, Tairan Huang, Yichao Cao, Hongyan Xu, Keke Huang, Haifeng Li, Yi Chen, Xiu Su

この論文をやさしく読む

ひとことで言うと

火災で相互に関わる複数の物理現象を使い、AIが見えていない状態や介入後の変化を推論できるか測る評価データ。

何に役立つ?

マルチモーダルなモデルやエージェントの物理的な推論能力を、火災の場面で比較・分析するために使える。

この研究の面白いところ

494件の制御されたシミュレーションと26件の実世界に対応する事象群から、文章・2D物理場・3D場面を組み合わせた9074組の質問回答を作った。

どこまで分かった?

要旨は評価基盤の構成と測定対象を示すが、個別のモデルの得点や、実際の火災現場での運用性能は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

物理世界を理解するには、物体認識、場面の説明、短期的な映像予測だけでは足りない。実際の物理系には、複数の連続的な場、隠れた因果的仕組み、部分的な観測、介入で変わる動きがある。本研究では、相互に結合した火災の物理場を通じて、マルチモーダルな大規模言語モデルやエージェントの複雑な物理世界の理解を評価するFireWorldBenchを提案する。火災は、複数の物理場が観測可能な状態と時間変化を共同で作る、負荷の高い代表的な環境である。評価基盤は、物理能力と火災場面の課題という二つの軸で構成され、物理状態の理解、時間的な動き、因果的な仕組み、介入についての推論を扱う。 47種類の場面の基本型と7種類の環境群にまたがる520件の火災世界の項目を含み、その内訳は制御されたシミュレーション世界494件と実世界に対応付けた事象群26件である。項目は、構造化された文章の観測、複数の2次元物理場の画像、事象レベルの3次元場面模型を組み合わせる。選択式の問題と自由記述の報告生成を合わせ、文章と画像が交互に現れる9074組の質問・回答を用意した。部分的なマルチモーダル観測から、隠れた物理状態を推測し、仕組みを説明し、結合した場の変化を予測し、介入の結果を評価できるかを測るための、難しい試験基盤となる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Understanding the physical world requires more than object recognition, scene description, and short-term visual prediction, as real-world physical systems involve multiple continuous fields, latent causal mechanisms, partial observations, and intervention-sensitive dynamics. We propose FireWorldBench, a benchmark for evaluating complex physical world intelligence in multimodal large language models and agents through coupled-field fire dynamics. Fire provides a canonical stress-test environment, where multiple interacting physical fields jointly shape observable states and temporal dynamics. FireWorldBench is organized along two complementary axes, a physical capability axis and a fire scenario task axis, jointly covering physical-state understanding, temporal dynamics, causal mechanisms, and intervention reasoning. The benchmark comprises 520 fire-world entries, including 494 controlled simulation worlds and 26 real-world-aligned event groups, spanning 47 scene archetypes across 7 environment families. These entries combine structured textual observations, multiple 2D physical-field visualizations, and 3D event-level scene modeling, yielding 9,074 text-image interleaved question-answer pairs across choice-based and open-ended report-generation formats. FireWorldBench evaluates whether models can infer latent physical states, explain underlying mechanisms, forecast coupled-field evolution, and assess intervention consequences from multimodal partial observations, providing a challenging testbed for complex physical world intelligence.

arXiv ID: 2609.23064 / 要約の誤りについて