障害物との絡み方を画像上の目印で伝えるロボット操作
Topology-Informed Visual Prompting For Vision Language Action Policies
この論文をやさしく読む
ひとことで言うと
障害物に対して腕や物体がどう絡んでいるかを推定し、次に進む位置をカメラ画像へ描き込んでロボットを誘導します。
何に役立つ?
見た目が似ていても回り込む方向などが異なる操作で、ロボットに必要な経路の手がかりを渡す用途があります。実機では箱の持ち上げを検証しています。
この研究の面白いところ
学習時はシミュレーションの詳しい形状情報を使い、運用時はカメラ画像だけから特徴量と経由点を推定します。VLAへの指示も画像上の目印として渡します。
どこまで分かった?
評価はシミュレーション3課題と実機1課題です。実機の40%改善について、要旨は相対増加かパーセントポイント差かを明示していないため、そのどちらかと断定できません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚・言語・行動(VLA)方策は、部分観測のため、複雑な障害物形状を持つ操作課題で苦戦する場合がある。複雑な形状では、似た視覚観測やロボット配置であっても、質的に異なる行動が必要になり得る。この違いは位相的な特徴量で定量化できる。環境形状と物体状態を完全に知る運動計画器なら、その特徴量を考慮して計画できるが、実運用時にはそうした情報が分からないことが多い。 この問題に対し、シミュレーションに基づく計画で基本のデモンストレーションデータを拡張し、実運用時には視覚に基づく誘導を行う、位相に導かれた視覚プロンプトの枠組みを提示する。Gauss絡み数積分による位相的特徴量の表現を使い、環境の重要な位相的性質を捉える。環境の近似シミュレーションから得た特権的な幾何情報を用い、デモで示された特徴量の状態へ系を移してから、その新しい配置で課題を再開する軌道を、VLAの微調整データに追加する。 視覚言語モデル(VLM)も同じデータで微調整し、リアルタイムのカメラ観測から位相的特徴量とエンドエフェクタの経由点を予測させる。経由点は観測画像上に視覚プロンプトとして描画され、VLAを誘導する。シミュレーションの両腕操作3課題と、実世界の箱の持ち上げ課題で、基本デモだけで微調整したVLA、および観測から位相的に重要な情報を除いてしまう場合があるVLMプロンプトのベースラインを上回った。実機では、課題成功の指標で最も強いベースラインを40%上回った。プロジェクトのウェブサイトはhttps://topology-vla.github.ioである。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language-action (VLA) policies can struggle with manipulation tasks with complex obstacle geometries due to partial observability. These complex geometries can lead to similar visual observations or robot configurations requiring qualitatively different actions, a distinction that can be quantified using topological signatures. While motion planners with full knowledge of environment geometries and object states can reason about these signatures in planning, this information is often not known at deployment. To address this issue, we present a topology-guided visual-prompting framework that uses simulation-based planning to augment a nominal demonstration dataset and provides vision-based guidance at deployment. Our method uses a Gauss-Linking-Integral topological signature representation to capture important topological properties of the environment. Using privileged geometry information from a simulation approximation of our environment, we augment a VLA fine-tuning dataset with trajectories that move the system to a demonstrated signature and, from the new configuration, resume task execution. A vision-language model (VLM) is fine-tuned on the same dataset to both predict signatures from live camera observations and predict end-effector waypoints, which are rendered as visual prompts on the observations to guide the VLA. Across three simulated bimanual tasks and a real-world box pickup task, our method outperforms a VLA fine-tuned only on nominal demonstrations and a VLM-prompting baseline that can remove topology-relevant information from observations. On hardware, it exceeds the strongest baseline by 40% in task success. Project website: https://topology-vla.github.io.
arXiv ID: 2609.23944 / 要約の誤りについて