操作に必要な画像領域へロボットの視線を誘導
ActGaze: Learning Action-Grounded Gaze through Counterfactual Visual Interventions for High-Precision Manipulation
この論文をやさしく読む
ひとことで言うと
ロボットが精密な作業をするとき、行動に必要な画像の領域へ注意を集中させる訓練方法。
何に役立つ?
視覚・言語・行動モデルを使った精密操作の改善に役立つ可能性がある。
この研究の面白いところ
人が視線位置にラベルを付ける代わりに、画像を反事実的に変えて行動予測への影響を調べ、注目領域を決める。
どこまで分かった?
実証は要旨に記された4つの高精度操作課題での実機実験に限られ、改善量の具体的な数値は示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
現在の視覚・言語・行動モデル(VLA)は、高精度なロボット操作を苦手とすることが多い。著者らは主な原因を、視覚的注意が課題に関係しない領域へ分散することだと考える。この問題に対し、精密な動作中に人が重要な視覚的手掛かりを見るのと同様、VLA 方策の視線を課題関連領域へ向ける訓練法 ActGaze を提案する。視線の教師信号に外部のラベルを用いる従来法とは異なり、反事実的な視覚介入を使って行動予測に重要な領域を特定し、VLA 自身の行動目標から空間的な教師信号を得る。4つの高精度ロボット操作課題で行った広範な実機実験では、ActGaze によって課題関連領域への視覚的注意がより集中し、元の VLA 方策と他の視覚的な位置付け手法を一貫して上回った。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Current Vision-Language-Action (VLA) models often struggle with high-precision robotic manipulation. We attribute this limitation primarily to their visual attention being dispersed across task-irrelevant regions. To address this issue, we propose ActGaze, a training approach that guides VLA policies to gaze on task-relevant regions, much like humans gaze on critical visual cues while executing precise movements. Unlike prior methods that rely on external labels for gaze supervision, ActGaze derives spatial supervision directly from the VLA's own action objective by using counterfactual visual interventions to identify regions that are critical for action prediction. Extensive real-robot experiments on four high-precision robotic manipulation tasks demonstrate that ActGaze induces more focused visual attention on task-relevant regions and consistently outperforms the base VLA policy and other visual-grounding approaches.
著者のコメント
11 pages, 7 figures
arXiv ID: 2609.28955 / 要約の誤りについて