arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

言語の空間指示をLiDARの座標に結び付ける

Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

Byounggun Park, Giyong Moon, Jusung Kim and Soonmin Hwang

この論文をやさしく読む

ひとことで言うと

複雑な位置関係を含む言葉の質問から、指された物体の座標をLiDAR点群内で求める研究です。

何に役立つ?

考えられる用途は、屋外ロボットや自動運転での言語による対象指定です。要旨で実証したのは座標予測課題での比較性能です。

この研究の面白いところ

言語モデルで質問を解釈しつつ、最終的な座標は文章として推測せず、対象付近のLiDAR点群から直接求めます。

どこまで分かった?

要旨は代表的な比較手法より良い結果を述べていますが、改善幅の数値や実環境での運用結果は記していません。データと訓練コードは公開予定です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

LiDARは、自動運転や屋外ロボットの物体検出などで、正確な幾何情報を与える。しかし、個々の物体を認識して位置を特定するだけでは、複数の空間的関係を組み合わせ、質問が指す対象を特定する必要がある問いには答えられない。自動運転向け大規模言語モデルの進展を踏まえ、本研究は言語モデルの知識を使って複雑な空間質問を解釈し、指示対象をLiDARの幾何情報に結び付ける。 この能力を支えるため、単段階と多段階の関係に基づく対象特定と、補完的な空間理解課題を組み合わせたSpatialLiDAR-QAを導入する。また、LiDAR点群の特徴を大規模言語モデルに対応付け、言語で条件付けた位置情報を考慮する候補検索と、局所点群での詳細化によって対象座標を特定するSpatialLiDAR-LMを提案する。この設計では、言語として座標を生成するのではなく、局所的なLiDAR幾何情報から直接座標を求める。 実験では、正確な座標予測の課題で、代表的なLiDARと言語を組み合わせたモデル、および複数カメラを使う視覚言語モデルを大幅に上回った。データセットとモデルの訓練コードは公開予定である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

LiDAR provides precise geometric information for spatial perception tasks such as object detection in autonomous driving and outdoor robotics. However, recognizing and localizing individual objects is not sufficient to answer questions that require composing spatial relations and grounding the intended target. Motivated by recent advances in large language models (LLMs) for autonomous driving, we leverage their language priors to interpret complex spatial questions and ground the referred target in LiDAR geometry. To support this spatial grounding capability, we introduce SpatialLiDAR-QA, which combines single- and multi-step relational grounding with complementary spatial understanding tasks. We further propose SpatialLiDAR-LM, which aligns LiDAR point features with an LLM and grounds target coordinates through language-conditioned, position-aware proposal retrieval and local point refinement. This design derives target coordinates directly from local LiDAR geometry rather than through textual language decoding. Experiments demonstrate substantial improvements over representative LiDAR--language models and multi-camera VLMs on precise coordinate prediction tasks. Our dataset and model training code will be publicly released.

著者のコメント

8 pages

arXiv ID: 2609.29835 / 要約の誤りについて