arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

未来の動きと視覚構造を学びロボットの行動生成を改善

MT-WAM: Reorienting the One-Pass Predictive Representation Toward Action Generation

Yiguang Yang, Jiankun Peng, Xiaoming Wang, Yiran Zhang, Zhibo Fang

この論文をやさしく読む

ひとことで言うと

ロボットが次の行動を決めやすいよう、映像の未来予測に加え、物の動きと見た目の構造も学ばせる方法です。

何に役立つ?

視覚条件が変わる場面での行動生成を改善する手掛かりになります。ベンチマークに加え、四つの実世界タスクでの成功率も報告されています。

この研究の面白いところ

学習時には未来の情報を複数の形で教えますが、実行時は未来動画を生成しません。また、動きの特徴は行動へ渡し、視覚特徴は教師信号として使うという役割分担があります。

どこまで分かった?

LIBEROなどのベンチマークと実世界4タスクの結果は別の評価です。23.8は相対改善率ではなくパーセントポイント差で、要旨には実世界タスクの内訳や試行数はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Fast-WAMは、推論時に未来の動画を生成しなくても動画と行動の共同学習によって制御が改善することを示しており、動画拡散Transformerの1回の順伝播で得る表現が、行動生成の中心になる。しかし、未来の観測を予測することは、制御に必要な未来の動力学や視覚構造を明示的に優先するわけではない。 本研究では、元の学習目的を保ちつつ、未来の2次元点軌道と視覚特徴について補完的な教師信号を加えるMT-WAMを示す。動画バックボーンの最後のブロック群をコピーした軽量な2ストリーム分岐が、目標に特化した処理を行い、構造化した注意マスクがストリーム間の注意を防ぐ。動きのストリームのトークンは、行動エキスパートに動力学の追加条件を与える。未来の視覚特徴の予測は、物体と空間構造を捉える特徴空間で教師信号を提供する。この教師信号によって、視覚条件が変化する下でも行動生成に有益な視覚的文脈を提供するよう動画バックボーンを学習させるが、視覚特徴ストリームのトークンを行動の条件付けに追加することはしない。推論時には、再計画のたびに一度だけ計算する動画と動きのキャッシュを使い、未来動画の予測は省略する。 身体を持つエージェントの方策に対する追加の事前学習なしで、MT-WAMはLIBEROで98.2%、LIBERO-Plusで73.7%の成功率を達成し、後者ではFast-WAMを23.8パーセントポイント上回る。RoboTwin 2.0 Clean2Randでは、Random条件の成功率が6.30%から19.40%へ上昇する。四つの実世界タスクでは、平均成功率が67.0%から77.8%へ上昇する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Fast-WAM shows that video-action co-training improves control without generating future video at inference, making the representation from a single video diffusion Transformer forward central to action generation. However, future-observation prediction does not explicitly prioritize the future dynamics and visual structure needed for control. We present MT-WAM, which retains the original training objectives and adds complementary supervision for future two-dimensional point trajectories and visual features. A lightweight dual-stream branch copied from the video backbone's final blocks provides target-specific processing, while a structured attention mask prevents cross-stream attention. Motion-stream tokens supply additional dynamics conditions to the action expert. Future visual-feature prediction provides supervision in a feature space that captures object and spatial structure. This supervision trains the video backbone to provide more informative visual context for action generation under changing visual conditions, without adding visual-feature-stream tokens to action conditioning. At inference, MT-WAM uses video and motion caches computed once per replan and skips future-video prediction. Without additional embodied policy pretraining, MT-WAM achieves 98.2% success on LIBERO and 73.7% on LIBERO-Plus, exceeding Fast-WAM by 23.8 percentage points on the latter. On RoboTwin 2.0 Clean2Rand, Random success increases from 6.30% to 19.40%; across four real-world tasks, average success increases from 67.0% to 77.8%.

arXiv ID: 2609.21474 / 要約の誤りについて