遮蔽物の不確かさを考慮して物をつかむCPOR-Grasp
Calibrated Probabilistic Obstruction Reasoning with Vision-Language Models for Grasping in Clutter
この論文をやさしく読む
ひとことで言うと
物同士の隠れ方に複数の解釈を残し、直接つかむか、邪魔な物を動かすか、保留するかを確率に基づいて決めるロボット向けの方法です。
何に役立つ?
散らかった場面で、単一の画像解釈に決め打ちした誤判断を減らすことに役立ちます。計算量を抑えながら保留判断も扱えます。
この研究の面白いところ
物体対ごとの予測を整合的なグラフにまとめ、計算のために捨てるグラフが判断に与える影響まで評価しています。
どこまで分かった?
99.74%は近似判断と厳密推論の一致率で、現実の把持成功率ではありません。実環境の平均成功率は77.8%です。打ち切り誤差の保証も、視覚認識を含む全行動の安全性の保証とは別です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
物が入り混じる場所から目標物を取り出すには、目標をつかむか、遮る物を除くか、判断を保留するかを決める必要がある。既存手法は通常、単一の遮蔽グラフや除去戦略に決め打ちし、複数の場面解釈の間にある不確かさを無視する。また、較正の不十分な視覚言語モデル(VLM)の予測に依存し、対ごとの遮蔽関係が全体では矛盾することもある。さらに現在の近似には、捨てた仮説が最終判断に与える影響について保証がない。 そこで、対ごとの証拠から行動判断まで不確かさを伝播させる、較正された確率的遮蔽推論の枠組みCPOR-Graspを提案する。VLM、深度、物体の隠れた部分も含むアモーダルマスクの手掛かりを較正して統合し、遮蔽確率を推定する。矛盾のない遮蔽グラフ上の分布を作り、それらについて周辺化することで、目標に接近可能である確率や、特定の遮蔽物を除くべき確率を計算する。推論を計算可能にするため、高確率のグラフだけを保持し、捨てた確率質量に関する全変動距離の上界を導く。これにより、保証付きの判断、適応的な打ち切り、原則に基づく判断保留が可能になる。 合成および実際のUNOBench場面で、CPOR-Graspは最先端の比較手法を上回る。Gemini Roboticsを基盤モデルにした場合、較正誤差は0.1416から0.0185へ低下する。グラフを打ち切る近似は、使用するグラフ数を56分の1にしながら、判断の99.74%で厳密推論と一致する。実環境の実験では平均成功率77.8%を達成し、最先端の比較手法を上回る。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Retrieving a target from clutter requires deciding whether to grasp the target, remove a blocker, or defer. Existing methods typically commit to a single obstruction graph or removal strategy, ignoring uncertainty across alternative scene interpretations. They also rely on miscalibrated vision-language model (VLM) predictions and can produce pairwise obstruction relations that are jointly inconsistent. Moreover, current approximations provide no guarantees about the impact of discarded hypotheses on the final decision. We propose CPOR-Grasp, a calibrated probabilistic obstruction-reasoning framework that propagates uncertainty from pairwise evidence to action decisions. CPOR-Grasp calibrates and fuses VLM, depth, and amodal-mask cues to estimate obstruction probabilities, induces a distribution over valid obstruction graphs, and marginalizes over these graphs to compute the likelihood that the target is accessible or that a given blocker should be removed. To make inference tractable, it retains only the highest-probability graphs and derives a total-variation bound on the discarded probability mass, enabling certified decisions, adaptive stopping, and principled deferral. On synthetic and real UNOBench scenes, CPOR-Grasp outperforms state-of-the-art baselines. Calibration error decreases from 0.1416 to 0.0185 on the Gemini Robotics backbone, while graph truncation matches exact inference on 99.74\% of decisions using 56 times fewer graphs. In real-world experiments, CPOR-Grasp achieves a 77.8\% average success rate, surpassing SOTA baselines.
著者のコメント
Submitted to ICRA
arXiv ID: 2609.18718 / 要約の誤りについて