物理世界の理解・映像生成・ロボット動作をまとめて学習
UniWAM: Unified World-Action Model
この論文をやさしく読む
ひとことで言うと
映像と言葉の理解、将来の見え方の生成、ロボットの行動予測を一つのモデルで学習する方法です。人間の一人称映像とロボットの実演を組み合わせます。
何に役立つ?
考えられる用途は、言葉の指示を理解して長い手順の作業を行うロボットです。複数の能力評価で良好な性能と、生成時のノイズ除去ステップの削減が報告されています。
この研究の面白いところ
低水準の動作を自然言語で表し、異なる種類のデータを各構成要素へ割り当てています。将来映像への雑音追加と行動履歴の利用で、将来予測への依存と計算量を調整しています。
どこまで分かった?
要旨には評価タスク、成功率、データ量、削減したステップ数の具体値がありません。対数線形のスケーリング則が確認された範囲も記載されていないため、任意の規模への外挿はできません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚・言語・行動モデルは、事前学習済みの視覚言語モデルの理解力と推論能力を活用できるが、行動だけを教師情報とする学習では、世界の動力学への基礎づけが限られる。一方、世界・行動モデルは動画生成モデルの時空間的な事前知識を受け継ぐが、分布が変わる状況では意味理解や推論に限界がある。本研究では、物理的な推論器、世界生成器、行動予測器を統合し、物理世界の意味理解、視覚生成、行動予測を共同で学習する統一アーキテクチャUniWAMを導入する。 学習データの品質を確保するため、人間の一人称視点データとロボットデータの両方に対し、厳密なデータ整理・注釈付与の処理手順を開発した。受け継いだ言語能力を保ちながら視覚言語部分を身体を伴うタスクに適応させるため、低水準の動作を自然言語で表現する。また、視覚質問応答(VQA)データ、人間の一人称視点データ、ロボットの実演から得られる相補的な教師情報を適切なモデル構成要素に割り当てる事前学習手順を導入する。事後学習では、将来の視覚情報に雑音を加える拡張により、正確な将来予測への依存を減らす。同時に、履歴を条件とするフローマッチングでは、符号化した行動履歴を使って行動生成を初期化する。これらの設計を組み合わせることで、性能を保ちながらノイズ除去のステップ数を大幅に減らす。 UniWAMは、分布内性能、頑健性、一般化、指示への追従、長期にわたるタスクの実行を含む複数の評価で、最先端の性能を達成する。さらに、人間とロボットを統合した共同学習に対数線形のスケーリング則があることを見いだし、人間とロボットの混合データを用いた大規模事前学習の有効性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predictor to jointly learn semantic understanding of the physical world, visual generation, and action prediction. To ensure the quality of the training data, we developed a rigorous data cleaning and annotation pipeline for both human egocentric data and robot data. To adapt the vision-language component to embodied tasks while preserving its inherited language capabilities, we represent low-level actions in natural language and introduce a pre-training recipe that assigns complementary supervision from visual question answering (VQA) data, human egocentric data, and robot demonstrations to the appropriate model components. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while history-conditioned flow matching uses encoded action history to initialize action generation. Together, these designs significantly reduce denoising steps while maintaining performance. UniWAM achieves state-of-the-art (SOTA) performance across multiple evaluations, including in-distribution performance, robustness, generalization, instruction following, and long-horizon task execution. Furthermore, we uncover a log-linear scaling law of unified human-robot co-training, demonstrating the effectiveness of large-scale pre-training on a mixture of human and robot data.
arXiv ID: 2610.02054 / 要約の誤りについて