行動に必要な未来情報だけを残すロボットの予測状態
The Right Future for Action: Learning Action-Relevant Predictive States in World Action Models
この論文をやさしく読む
ひとことで言うと
ロボットの行動を決めるとき、未来の映像を全部作らず、行動に役立つ未来情報を小さな状態にまとめる。
何に役立つ?
視覚から行動を選ぶロボット方策の、環境が変わったときの一般化に役立つ。
この研究の面白いところ
未来の変化を予測しやすいだけでは動作を読み出しやすくならないという診断から、行動用の状態表現を設計した。
どこまで分かった?
成績はLIBEROとLIBERO-Plusの指定評価であり、実機での成功率は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
画像を生成しない世界・行動モデル(WAM)は、学習時に未来の動画を予測するが、推論時には内部の動画特徴から動作を選ぶ。そこで、その特徴が制御のために何を保つべきかを調べる。表現の診断では、未来の変化を予測しやすい表現が、必ずしも動作を線形に読み出しやすくするわけではなかった。観測された未来の変化は現在の情報に加えて動作に関する情報を持ち、線形に読み出せる動作情報は空間的に集中している。この知見から、動画側と行動側の間に置く小さな予測表現Action-Relevant Predictive States(ARPS)を提案する。予測する時間幅に条件付けた状態予測器が、中間の動画特徴を、行動側に渡す全視覚情報を含む小さな状態へ集約する。未来の表現を使う教師信号により、その状態の異なる部分が別々の未来時点の視覚表現と、現在からの変化を予測するよう学習する。推論時には教師信号の分枝を外し、現在の観測から作る予測状態だけを行動側が使う。要素を分けた評価では、未来の教師信号が分布変化のもとでの一般化を大きく改善した。ARPSはLIBEROで成功率99.2%、追加の適応なしのLIBERO-Plusで87.3%に達し、Fast-WAMより39.2パーセントポイント高かった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Generation-free world action models (WAMs) retain future-video prediction during training but act from internal video features at inference, leaving unclear what these features should preserve for control. Our representation diagnostics show that representations with more predictable future changes need not make linear action decoding easier. Observed future changes provide additional action information beyond the present, and linearly readable action information is spatially concentrated. These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts. ARPS uses a horizon-conditioned state predictor to aggregate intermediate video features into a compact state that supplies all visual context to the action expert. Future-representation supervision trains different parts of this state to predict visual representations at different future times, together with their changes relative to the present. At inference, the supervision branch is removed, and the action expert only uses the learned predictive state computed from current observations. Controlled ablations show that future supervision substantially improves generalization under distribution shift. ARPS achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.
著者のコメント
19 pages, 5 figures, 8 tables
arXiv ID: 2609.23369 / 要約の誤りについて