全方向視覚を使う脚ロボットの姿勢認識探索
Pose-aware Legged Robot Semantic Exploration with Omnidirectional Perception in Confined Unknown Environments
この論文をやさしく読む
ひとことで言うと
脚ロボットが体を傾ける姿勢まで計画し、狭い場所で物体の見えにくい面を効率よく観察する仕組みです。観察済みの履歴も使って重複した訪問を減らします。
何に役立つ?
設備や物体の探索・点検が考えられる用途です。工作機械のある作業場で実機実験を行い、適用可能性を示しています。
この研究の面白いところ
視野を広げる姿勢変更と、その動作にかかる時間を同時に考えています。視覚言語モデルによる視点の絞り込みと全体の探索計画を組み合わせています。
どこまで分かった?
面の観察率8~10ポイント改善、時間17~32%短縮という数値はシミュレーションの平面計画との比較です。同じ数値が実機でも得られたとは要旨に書かれていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
狭い未知環境での意味探索には、環境の地図作成と対象物の詳細な観察の両方が必要です。地上ロボットでは、センサーの垂直視野が限られ、対象物から離れて観察できる距離も制約されるため、平面的な視点からは対象物の上面が見えないことがあります。機体を傾ければ被覆範囲を広げられますが、追加観測と姿勢遷移によってミッション時間が増えます。 このトレードオフに対処するため、脚ロボット本来の機体のピッチとロール、および全方向カメラ・LiDAR認識を利用する、姿勢認識型意味探索システムPOSEを提案します。姿勢認識型の視点サンプリングモジュールは、部分的な対象物地図から期待される被覆向上に基づいて機体姿勢を選び、照準に合わせた実行によって不要な機体の向き直しを減らします。さらに、永続的な観測履歴と鳥瞰図(BEV)地図を使い、視覚言語モデル(VLM)の支援で物体中心の視点を刈り込む戦略を導入して、重複した訪問を減らしました。得られた意味視点を幾何探索視点と組み合わせ、全体探索プランナーに統合します。 シミュレーションでは、POSEは平面計画を基準に、最終的な対象表面の被覆率を8~10パーセントポイント高め、探索時間を17~32%短縮しました。また、評価した基準手法の中で物体被覆AUCの平均が最も高くなりました。全方向カメラ・LiDAR一式を搭載した脚ロボットを機械工場で動かす実世界実験でも、システムの適用可能性を示しました。これらの結果は、脚ロボットの意味探索で被覆と効率のトレードオフを改善するには、適応的な機体姿勢計画が有効であることを支持します。コードは将来、コミュニティのために公開する予定です。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Semantic exploration in confined environments requires both environment mapping and detailed observation of target objects. For ground robots, limited sensor vertical fields of view and restricted standoff distances can leave upper object surfaces unobserved from planar viewpoints. Body tilting can improve coverage, but additional observations and posture transitions increase mission time. To address this trade-off, we present POSE, a pose-aware semantic exploration system that exploits a legged robot's intrinsic body pitch and roll with omnidirectional camera-LiDAR perception. The proposed pose-aware viewpoint sampling module selects body postures from partial object maps according to expected coverage gain, while aim-aligned execution reduces unnecessary body reorientation. Further, we introduce an object-centric viewpoint pruning strategy assisted by a vision-language model (VLM), which uses persistent observation history and bird's-eye-view (BEV) maps to reduce redundant inspection visits. The resulting semantic viewpoints are combined with geometric exploration viewpoints in a global exploration planner. Simulations show that POSE improves final target-surface coverage by 8-10 percentage points over the planar planning baseline while reducing exploration time by 17-32%, and achieves the highest mean object coverage AUC among the evaluated baselines. Real-world experiments with a legged robot carrying an omnidirectional camera-LiDAR suite in a machine shop further demonstrate the system's applicability. These results support adaptive body-posture planning for improving the coverage-efficiency trade-off in legged robot semantic exploration. We plan to release the code for community benefit in the future.
著者のコメント
Submitted to ICRA 2027 under review
arXiv ID: 2609.19460 / 要約の誤りについて