未来の観測予測をロボットの行動生成につなぐATI-VLA
ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection
この論文をやさしく読む
ひとことで言うと
将来の映像などを予測するモデルの情報を、ロボットの行動決定で使いやすい形へ変換して取り込む方法です。
何に役立つ?
予測能力を持たせても操作性能が上がらない問題に対し、予測と行動の表現をそろえて活用する設計の参考になります。
この研究の面白いところ
予測と行動を共有の離散表現に合わせてから、行動生成へ補助経路で注入します。複数目的の競合を避ける行動中心の構成を採ります。
どこまで分かった?
シミュレーションと実機の双方で評価されていますが、要旨にはタスク数、成功率、収束速度の具体的な値はありません。最先端性能という主張の対象範囲は評価されたタスクです。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
予測型の視覚・言語・行動(VLA)モデルは、将来の観測や世界の力学を予測して、ロボット操作を改善することを目指す。しかし、既存手法はその可能性を十分に実現できず、行動を直接予測するモデルを下回ることが多い。本研究では、その制約が、観測と行動のモダリティ間の不整合、および学習を行動中心の目的から遠ざける同時最適化の競合に由来すると論じる。 そこで、行動に使える表現の整合を行ってから適応的に情報を注入する、行動中心の予測型VLAの枠組みATI-VLAを導入する。具体的には二段階で設計する。第一は、共有コードブックによる、行動に利用可能な表現の整合である。統一したコードブックを介して予測観測と行動の両方を共通の離散潜在空間へ写像し、表現をそろえる。これにより、予測観測の潜在表現を行動生成へそのまま使いやすくし、モダリティ間の不整合を軽減する。 第二は、行動を中心とした予測潜在表現の適応的な注入である。第一段階を基に、軽量な適応型の補助経路を通じ、予測観測の潜在表現を明示的な予測事前情報として行動の復号へ注入する。これにより、一つの行動中心の目的の下で、予測による誘導を適応的に行う。シミュレーションと現実のロボットタスクの両方における広範な実験で、ATI-VLAがより速く収束しながら最先端の性能を達成することを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment between observations and actions, together with joint optimization conflicts that drive learning away from an action-centric objective. To this end, we introduce ATI-VLA, an Action-Centric Predictive Vision-Language-Action framework via Actionable Alignment Then Adaptive Injection. Specifically, it follows a two-step design: 1) Actionable Representation Alignment via a Shared Codebook. It aligns predictive observation and action representations by mapping both modalities into a shared discrete latent space via a unified codebook, making predictive observation latents readily usable for action generation and mitigating modality misalignment. 2) Action-Centric Adaptive Injection of Predictive Latents. Building upon this, it then injects predictive observation latents into action decoding as explicit predictive priors via a lightweight adaptive side-path, enabling adaptive predictive guidance under a single action-centric objective. Extensive experiments on both simulation and real-world robotic tasks demonstrate that ATI-VLA achieves state-of-the-art performance with faster convergence.
著者のコメント
Accepted to NeurIPS 2026. Project page: https://jiutian-vl.github.io/ATI-VLA-page/
arXiv ID: 2610.01741 / 要約の誤りについて