arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

見えない物体を想定してロボットの観測と操作を計画する

Imagine-TAMP: Imagination-Guided Task and Motion Planning in Partial Observability

Antareep Singha, Shivaram Kumar, Yoonwoo Kim, and Yoonchang Sung

この論文をやさしく読む

ひとことで言うと

物陰の目標を探すロボットが、別の位置から見るか、邪魔な物を動かすかを事前に比べます。

何に役立つ?

棚など視界が限られる場所で、無駄な観測や操作、重い動作計画を減らす用途が考えられます。

この研究の面白いところ

意味的な位置の手掛かりと、隠れた場所の生成形状を使い、操作の労力と見つかりやすさを同時に見積もります。

どこまで分かった?

視点制約のある棚場面で成功率が46.0%から84.0%へ改善しました。実機では完全系の計画時間が幾何のみの比較版より32%短く、二つは異なる比較条件の結果です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

物が密集した環境で動作するロボットは、位置を部分的にしか観測できない物体を操作しなければならないことが多い。中心となる課題は、もう一度観測するか、先に目標を隠している可能性のある物体を動かすかの判断である。従来のタスク・動作計画TAMPは、通常、記号的な行動コストや計算負荷の高い幾何学的計画でこれを判断するが、どちらも観測によって隠れた目標が見える確率を十分に捉えられない。 Imagine-TAMPは、計画と実行を交互に進める枠組みであり、計算負荷の高い動作計画に着手する前に、意味的・幾何学的な想定を用いて、部分観測下の複数のタスク戦略を比較する。視覚言語モデルは、目標と見えている物体の常識的な関係を使って目標位置の粒子信念分布を形作り、生成シーンモデルは未観測領域のもっともらしい形状を推定する。 目標位置の仮説と想定シーンを与えると、Imagine-TAMPは複数の記号的な計画骨格を生成し、操作の労力と観測行動からの目標の見えやすさの両方を近似する、一律ではないコストを付ける。これにより、短いが情報の乏しい観測戦略と、先に遮蔽物を動かして目標を見えやすくする長い戦略を区別する。選択した骨格を実行可能な連続計画へ具体化して実行し、新しい観測で信念分布を更新して、必要なら再計画する。 実験では、想定に基づく評価によって観測と操作の選択が改善した。視点が制限された棚の場面では、一律ではない幾何学的評価により成功率が46.0%から84.0%へ上昇し、意味情報による信念分布の調整は操作と再計画をさらに減らした。実機ロボットでは、完全なシステムは幾何情報だけを用いるアブレーション構成と比べて計画時間を32%短縮した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robots operating in cluttered environments must often manipulate objects whose locations are only partially observable. A central challenge is deciding whether to acquire another observation or to first manipulate objects that may occlude the target. Conventional task and motion planning (TAMP) approaches typically make this decision using symbolic action costs or expensive geometric planning, neither of which adequately captures how likely an observation is to reveal an occluded target. We introduce Imagine-TAMP, an interleaved planning and execution framework that uses semantic and geometric imagination to compare alternative task-level strategies under partial observability before committing to expensive motion planning. A vision-language model shapes a particle belief over target locations using commonsense relationships between the target and visible objects, while a generative scene model estimates plausible geometry in unobserved regions. Given a target hypothesis and imagined scene, Imagine-TAMP generates multiple symbolic plan skeletons and assigns non-unit costs that approximate both manipulation effort and target visibility from sensing actions, distinguishing a short but poorly informative observation strategy from a longer strategy that first manipulates an occluder to better expose the target. The selected skeleton is then refined into a feasible continuous plan and executed, with new observations updating the belief and triggering replanning when necessary. Experiments show that imagination-guided evaluation improves observation-versus-manipulation decisions: in viewpoint-constrained shelf scenes, non-unit geometric evaluation increases success from 46.0% to 84.0%, while semantic belief shaping further reduces manipulation and replanning. On a real robot, the complete system reduces planning time by 32% relative to a geometry-only ablation.

arXiv ID: 2609.20396 / 要約の誤りについて