実環境で言葉による物体指定を評価するPro-Bench
Pro-Bench: Prompt-Robust Open-Vocabulary Visual Grounding Across Real-World Heterogeneous Environments
この論文をやさしく読む
ひとことで言うと
ロボットが言葉で指定された物体を実環境の画像から見つける能力を、多様な言い方で測るベンチマーク。
何に役立つ?
開いた語彙の視覚モデルを、実際のロボット環境や言い換えへの強さまで含めて比較できる。
この研究の面白いところ
16構成中10構成では短いカテゴリー名が最も良く、自由形式の問いが最良だったのは一構成だけだった。
どこまで分かった?
評価は収録された五種類の環境と16構成による。要旨は全モデルが実運用で十分頑健だとは主張していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
開いた語彙による視覚的な対象特定は、事前に決めた認識分類に頼らず、自然言語の問いからロボットが作業に関係する対象を見つけられるようにする。しかし既存のベンチマークは短いカテゴリー名とウェブから集めた画像に主に依存し、実運用の多様な問いや視覚条件でも頑健に対象を特定できるかは不明だった。本研究は、異種の実世界環境における開いた語彙の視覚的対象特定を評価する、プロンプト条件付きの Pro-Bench を導入する。地下、工業施設、屋内、屋外、都市という独立したロボット領域から集めた1万3千枚超の RGB フレーム、手作業による7万4500件の物体注釈、515種類の対象指定を含む。指定にはカテゴリー、属性、関係、使用可能性、状態、部分と全体、否定、組み合わせの意味を含める。16種類の開いた語彙のモデル構成を、追加学習なしの厳密なゼロショット条件で評価し、IoU のしきい値ごとの位置特定精度、端から端までの推論遅延、プロンプトによる性能変動、言い換えをまたいだ対象の一貫した再発見を測定した。結果、プロンプトへの頑健性はモデル構造に強く依存した。16構成中10構成は短いカテゴリー名で最高性能となり、自由形式の問いで最も高い精度になったのは一構成だけだった。また、集計された平均適合率 mAP が似ていても、言い換えをまたぐ一貫した対象の再発見には大きな違いが隠れる。Pro-Bench はこうした差を系統的に評価し、プロンプトに頑健な対象特定の研究を支える。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perceptual taxonomies. However, existing benchmarks largely rely on short category labels and web-scraped imagery, leaving it unclear whether open-vocabulary models can robustly ground diverse queries and visual conditions under real deployments. We introduce \textbf{Pro-Bench}, a prompt-conditioned benchmark for open-vocabulary visual grounding in heterogeneous, real-world environments. Pro-Bench includes $13k+$ RGB frames from independent robotic domains (subterranean, industrial, indoor, outdoor, urban), with $74.5k$ manual instance annotations and $515$ target queries covering categorical, attributive, relational, affordance, state, part-whole, negative, and compositional semantics. We benchmarked $16$ open-vocabulary model configurations in strict zero-shot inference, measuring localisation accuracy across IoU thresholds, end-to-end inference latency, prompt-induced performance variation, and target recovery consistency. Our results show that prompt-robustness is strongly architecture-dependent. Most model configurations ($10/16$) perform best with short category labels, whereas free-form queries yield the highest accuracy for only one. Moreover, similar aggregate mAP can conceal substantial differences in consistent target recovery across reformulations. Pro-Bench enables systematic evaluation of these gaps and supports prompt-robust visual grounding. Pro-Bench: https://pro-bench.github.io/.
arXiv ID: 2609.27076 / 要約の誤りについて