見えない物体を探すロボットの把持判断
Selective Commitment for Language-Guided Object Retrieval under Partial Observability
この論文をやさしく読む
ひとことで言うと
見えない対象を探しながら、いつ把持するかを決めるロボットの研究。
何に役立つ?
物体が隠れる環境で、観測を続けるか把持するかの判断を設計する参考になる。
この研究の面白いところ
対象が見えない場合や存在しない場合も仮説として保持し、失敗後に再観測した点。
どこまで分かった?
シミュレーションは25試行で、全機能版は簡略版を常に上回らなかった。実機では視点変更の失敗で誤った保留も起きた。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
一部しか見えない環境で、言葉で指示された物体を回収するには、さらに情報を集めるか、場面に働きかけるか、候補をつかむか、判断を控えるかを決める必要がある。本研究は、参照する容器との関係で指定された対象の回収について、これらの判断を調整する閉ループの枠組みを提案する。追跡中の物体、まだ観測していない対象、対象が存在しない可能性を仮説として、対象の同一性、容器との関係、存在の有無に関する統合的な信念を維持する。視点に応じた視覚言語モデルのカテゴリ観測で信念を更新し、適合的な予測に基づく把持資格とロボットの実行可能性で把持に踏み切るかを決める。有限の計画期間を持つ信念空間での計画により、情報を集める行動を選ぶ。五つの場面にわたるシミュレーションでは提案法が25試行中19回成功し、作業に合わせた最良の比較法の25試行中12回を上回った。また、評価した方策で唯一、全場面で少なくとも一回成功した。要素を外した比較では視点をまたぐ記憶が一部遮蔽下の成功率を上げたが、全機能を使う仕組みが簡略版を常に上回ったわけではない。実機試験では再観測の閉ループ動作と、意図的に起こした把持失敗からの自律的復旧を示した一方、視点変更を失敗させると誤って判断を保留した。結果は、部分的にしか見えない状況で情報収集と選択的な把持判断を統一的に扱える可能性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability.
arXiv ID: 2609.23131 / 要約の誤りについて