基準物体の寸法を使って室内の距離を推定する視覚言語モデル
Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes
この論文をやさしく読む
ひとことで言うと
画像内にある寸法の分かった物を手がかりに、視覚言語モデルへ室内の距離や大きさを推定させる方法。
何に役立つ?
ロボットの操作や移動で必要な寸法推定を研究する際のベンチマークと学習法になる。実機での安全な動作を証明した結果ではない。
この研究の面白いところ
カメラの内部パラメータを使わず、基準物体の既知寸法と数値で検証できる報酬を活用する。空間課題と一般課題の双方で改善を報告した。
どこまで分かった?
改善率は要旨に挙げた各ベンチマークの比較値で、全モデルや全室内環境への一般化は示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
距離や大きさを数値で推論する能力は、ロボット操作や自律移動などの実世界で動くAIに重要だが、視覚言語モデルには難しい。現在の空間推論は画素単位の厳密な教師信号に頼ることが多く、局所的な最適化で幅広いマルチモーダル推論能力が損なわれたり、破滅的忘却が起きたりし得る。本研究は、文脈情報を使う距離推論を促すベンチマークMetric-Benchを提案する。画像内に実寸が分かる基準物体を置くことで、カメラの内部パラメータがなくても、二次元から三次元への対応を暗黙に学べるようにする。さらに、構造化した指示と検証可能な数値報酬を用い、基準物体に根差した距離推論向けの強化微調整方法MetricReasonerを示す。Metric-Benchでの広範な実験では、既存手法や、より大きな専有モデルを43.1%上回った。下流の実世界行動に関わる評価でも、空間課題に特化した比較対象よりRoboSpatialの総合正解率で30.4%、ERQAで9.3%改善した。一般的なベンチマークでもV★Benchで15.9%、BLINKで88.9%の改善が一貫して見られ、提案する適応が必ずしも一般的な視覚言語能力を損なうわけではないことを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Metric reasoning is a critical and challenging task for Vision Language Models (VLMs), playing a pivotal role in embodied AI tasks such as robotic manipulation and autonomous navigation. However, current spatial reasoning remains bottlenecked by rigid pixel-level supervision; such localized optimization often compromises general multimodal intelligence, triggering performance degradation or catastrophic forgetting of broad reasoning capabilities. To address these limitations, we introduce Metric-Bench, a focused benchmark designed to guide metric-spatial reasoning using contextual information. By incorporating in-image reference objects with known physical dimensions, Metric-Bench guides models to implicitly learn the 2D-to-3D mapping without camera intrinsics. We further present MetricReasoner, a task-adapted reinforcement fine-tuning recipe for reference-grounded metric reasoning, using structured prompts and verifiable numerical rewards. Extensive experiments on Metric-Bench demonstrate that our approach significantly enhances spatial metric understanding, outperforming existing and even larger proprietary models by 43.1\%, while improving downstream embodied performance over a spatial-specialized counterpart by 30.4\% on RoboSpatial overall accuracy and 9.3\% on ERQA, and additionally delivering consistent gains on general benchmarks (15.9\% on V$\star$Bench, 88.9\% on BLINK), indicating that the proposed adaptation does not necessarily compromise general VLM capabilities.
著者のコメント
Accepted to ECCV
arXiv ID: 2609.25841 / 要約の誤りについて