物理世界との相互作用から視覚言語モデルの空間推論を学ぶ
Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World
この論文をやさしく読む
ひとことで言うと
物が動いたり視点が変わったりした後の位置関係を、視覚言語モデルが追い続けられるようにする学習方法です。
何に役立つ?
考えられる用途は、物理空間で行動するシステムの状態把握です。要旨では複数の視覚言語モデルと空間課題で改善を報告しています。
この研究の面白いところ
静止画像の質問だけでなく、観測、行動、次の観測の組を使います。局所的な変化から長い行動列への統合までを三段階で学ばせています。
どこまで分かった?
要旨には具体的な改善幅は記載されていません。評価は用いたシミュレーションと現実の軌跡、空間ベンチマークに基づきます。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
空間推論は、視覚言語モデル(VLM)が物理世界を理解し、そこで行動するために欠かせない。動的な環境では、物体の移動や視点の変化による局所的な状態遷移を捉え、それらを長い行動列にわたって統合して空間状態を更新し続ける必要がある。しかし既存のVLMは両方の能力に制約がある。現在の空間学習は主に物体の属性や位置関係に関する静的な質問に着目し、状態遷移を直接教える機会が少ない。これに対し、相互作用の軌跡は、直前の観測、行動、直後の観測を自然に結び付け、局所的な状態遷移を直接学習させられる。軌跡全体は連続する遷移の依存関係も示す。そこで、相互作用を通じて物理世界の状態遷移をVLMに学ばせる枠組みSpatial-Interactorを導入する。学習は、L1の受動的な世界状態の遷移、L2の能動的な自己状態の遷移、L3の長期的な相互作用軌跡という3段階のカリキュラムに分ける。各段階の目的に合わせた課題と、シミュレーションおよび現実の相互作用軌跡からなるLearning from Spatial Interactionデータセット(LSI-108K)を構築する。二段階の学習では、L1とL2に教師あり微調整を適用して局所遷移をモデル化する。次に、方策に沿った蒸留を用い、区間ごとの遷移の記述を与えられた教師側が、生徒側自身の方策で生成した思考過程を指導する。これにより、生徒はL3の長い軌跡で連続する遷移を統合することを学ぶ。複数のVLMと空間ベンチマークを用いた実験では、局所的な遷移のモデル化と長期的な統合の両方で一貫した改善が見られた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.
arXiv ID: 2609.23038 / 要約の誤りについて