arXiv論文メモ
新着一覧
cs.RO / cs.AI / cs.CV · 査読状況未確認

作業に関係する形状と未来の動きに注目するロボットAI

FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models

Zhiyuan Gao, Di Wen, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Schäfer, Kunyu Peng, Michael Beetz

この論文をやさしく読む

ひとことで言うと

ロボットが場面のすべてを同じように学ぶのではなく、今の小さな作業に関係する物の形と、これから起きる動きに注目するよう学習させます。

何に役立つ?

考えられる用途は、細かな位置関係や長い手順が重要な操作の改善です。学習に使う幾何・動態モデルを推論時には実行しないため、それらを毎回動かさず知識を利用できる構成です。

この研究の面白いところ

現在の形状を捉える知識蒸留と、将来の3次元的な変化を捉える表現を組み合わせます。場面全体の情報量を増やすだけでなく、作業との関連性で学ぶ領域を絞ります。

どこまで分かった?

要旨ではシミュレーションと実機の両方の改善を報告していますが、タスク数、成功率、計算時間の数値は示されていません。推論時に補助モデルが不要であることと、実測の高速化が示されたことは区別が必要です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

事前学習済み視覚言語モデルを基盤とする視覚・言語・行動(VLA)モデルは、多様なロボット操作タスクで高い性能を示している。しかし、現在の2次元観測を直接行動へ対応付けるVLAモデルは、空間的・時間的な理解が十分でないことが多く、精密な操作や長い手順の操作で性能が制限される。近年の方法は、場面全体への幾何学的な教師信号と将来状態の予測を通じてVLAモデルを強化する。しかし、これらの方法は場面の冗長な情報の影響を受け、現在の相互作用に関係する形状や動態の学習からモデルの注意をそらすことがある。 この問題に対し、小タスクに基づく幾何知識の蒸留と暗黙的な世界モデル化を組み合わせ、現在の空間構造と将来の相互作用の動態の表現を学ぶ枠組みFOCAL-VLAを提案する。幾何学習を現在の小タスクへ集中させるため、幾何の潜在表現を小タスクに関係する画像領域の特徴と整合させることで、VGGTからVLAモデルへ幾何知識を移す。現在の相互作用が将来3次元的にどう変化するかを捉えるため、実演の現在フレームと将来フレームから得たTrack4Worldの特徴を使って、暗黙的な世界モデル化を取り入れる。 この2つの相補的な表現が共同で行動生成を導き、推論時にはVGGTやTrack4Worldを動かす必要がない。実験では、FOCAL-VLAがシミュレーションのベンチマークと実世界の操作タスクの両方でベースラインを上回った。プロジェクトサイトは https://zhiyuan-gao.github.io/FOCAL-VLA/ である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in precise and long-horizon manipulation. Recent methods enhance VLA models through geometric supervision and future-state prediction across the entire scene. However, these methods can suffer from redundant scene information, distracting the model from learning the geometry and dynamics relevant to the current interaction. To address this issue, we propose FOCAL-VLA, a framework that combines subtask-guided geometry distillation with implicit world modeling to learn representations of current spatial structure and future interaction dynamics. To focus geometric learning on the current subtask, we transfer geometric knowledge from VGGT to the VLA model by aligning geometry latents with features from subtask-relevant image regions. To capture the future 3D evolution of the current interaction, we incorporate implicit world modeling using Track4World features from current and future demonstration frames. The two complementary representations jointly guide action generation without running VGGT or Track4World at inference time. Experiments show that FOCAL-VLA outperforms baselines on both simulation benchmarks and real-world manipulation tasks. Project website: https://zhiyuan-gao.github.io/FOCAL-VLA/.

arXiv ID: 2609.21228 / 要約の誤りについて