arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

物体の位置関係の説明から3次元地図の候補を検索

PosEviLoc: Position-Conditioned Spatial Evidence for Language-Based 3D Localization

Tianyi Shang, Yike Shi, Zhenyu Li

この論文をやさしく読む

ひとことで言うと

「ある物体がこちら側、別の物体が向こう側」といった説明を、仮に立つ位置ごとに確かめて地図の候補を選ぶ方法です。文章全体の似方だけでなく、複数の位置関係が同時に成り立つかを使います。

何に役立つ?

考えられる用途は、言葉による周囲の説明から点群地図の該当区域を探すことです。既存の検索結果を並べ直す部品としても評価されています。

この研究の面白いところ

説明中の方向を物体だけの属性とせず、仮の検索位置との関係として計算します。正解の姿勢を使わずに、各位置で説明の何割が成立するかという根拠を作ります。

どこまで分かった?

評価するのは対象位置を含むサブマップを選ぶ粗い位置推定です。17と16は相対的な改善率ではなくRecall@1のパーセントポイント差であり、精密な位置・姿勢の誤差を示す数値ではありません。要旨には速度やパラメータ数の具体値はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

言語に基づく3次元位置推定では、近くの物体とその空間関係の説明から、対象位置を含む点群サブマップを検索する。既存手法は通常、検索文とサブマップを大域的な記述子へ圧縮するため、物体単位の意味や、複数の説明の間にある空間的な整合性が不明瞭になる可能性がある。そこで、テキストから点群への粗い位置推定に向けて、検索位置を考慮するPosition-Conditioned Evidence Localization(PosEviLoc)を提案する。 PosEviLocは大域的な照合に頼らず、明示的な意味的・空間的根拠を使って各候補サブマップを評価する。方向を、物体位置と仮定した検索位置によって共同で決まる関係としてモデル化する。得られるQuery-Position Spatial Evidence Field(QSEF)は、各仮想位置で検索文中の説明のどれだけの割合が支持されるかを測る。根拠場の構築に正解の検索姿勢を使わず、複数の説明の一致を明示的に捉える。Multi-Level Evidence Readout(MER)がこの根拠をコンパクトな表現へ集約し、軽量なMLPが検索スコアへ変換する。 五つのベンチマークで、PosEviLocはRecall@1においてMNCLを平均17パーセントポイント上回った。差し替え可能な再順位付け器として用いた場合にも、MNCLを平均16パーセントポイント改善した。さらに、既存手法より追加パラメータ数を大幅に抑え、より高速な推論を達成した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Language-based 3D localization retrieves the point-cloud submap containing a target position from descriptions of nearby objects and their spatial relations. Existing methods typically compress queries and submaps into global descriptors, potentially obscuring object-level semantics and cross-description spatial coherence. We propose Position-Conditioned Evidence Localization (PosEviLoc), a query-position-aware framework for coarse text-to-point-cloud localization. Instead of relying on global matching, PosEviLoc evaluates each candidate submap using explicit semantic and spatial evidence. It models direction as a relation jointly determined by an object position and a hypothetical query position. The resulting Query-Position Spatial Evidence Field (QSEF) measures the fraction of query descriptions supported at each hypothetical position, explicitly capturing their agreement without using the ground-truth query pose to construct the evidence field. A Multi-Level Evidence Readout (MER) summarizes this evidence in a compact representation, which a lightweight MLP converts into a retrieval score. Across five benchmarks, PosEviLoc outperforms MNCL by an average of 17 percentage points in Recall@1. When used as a plug-and-play reranker, it improves MNCL by an average of 16 percentage points. Moreover, PosEviLoc introduces substantially fewer parameters and achieves faster inference speed than existing methods.

arXiv ID: 2609.23534 / 要約の誤りについて