ロボット操作で視覚言語モデルの説明と距離測定を分担
VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation
この論文をやさしく読む
ひとことで言うと
視覚言語モデルに物体の意味を担当させ、位置と距離は深度画像で測るロボット向けの場面表現。
何に役立つ?
ロボットが初めて見る場所で物を扱うとき、物体の説明と位置情報を両立させる設計に役立つ。
この研究の面白いところ
単一のRGB-D観測から物体単位の表現を作り、直接VLMに位置を答えさせる方法と151場面で比較した。
どこまで分かった?
評価は卓上の151場面に基づく。要旨は作業計画への組込みを述べるが、実行成功率の数値は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
未知の環境でロボットが作業するには、物の意味を理解すると同時に、信頼できる寸法・位置の情報が必要だ。視覚言語モデルは意味の理解には強いが、幾何学的な推定は信頼性が低い。本研究は、既存の手法を組み合わせ、視覚言語モデルを使うモジュール型の場面認識の枠組みを提案する。一枚のRGB-D観測から場面を物体単位の領域に分割し、視覚言語モデルで意味の注釈を付け、深度情報と対応させることで、特定の作業に依存しない物体中心の表現を構築する。卓上の151場面での実験では、この分担により高い意味理解の性能を保ちながら、視覚言語モデルに直接推定させる場合と比べて、位置特定と深度推定が大幅に改善した。得られた表現を、ロボットの実行のための作業計画の枠組みにも組み込んだ。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Robotic operation in previously unseen environments requires both semantic understanding and reliable metric information. While vision--language models (VLMs) provide strong semantic capabilities, their geometric estimates remain less reliable. In this paper, we propose a VLM-driven, modular perception framework for scene understanding using off-the-shelf approaches. Starting from a single RGB-D observation, the scene is segmented into object-level regions, annotated by a VLM, and grounded with depth information to construct a task-independent object-centric representation. Experiments on 151 tabletop scenes show that the proposed decomposition preserves strong semantic performance while substantially improving localization and depth estimation over direct VLM inference. The resulting representation is also integrated with a task-planning framework for robotic execution.
arXiv ID: 2609.28184 / 要約の誤りについて