arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

4次元ガウス表現でLiDARの疑似ラベルを作る

SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting

Nitya Nanvani, Andras Palffy, Holger Caesar

この論文をやさしく読む

ひとことで言うと

画像モデルの知識を使い、動く物体を含むLiDARデータに3次元の意味ラベルを自動で付ける方法です。

何に役立つ?

考えられる用途は、自動運転などの3次元認識用データ作成の負担を減らすことです。論文で示したのはSemanticKITTIでの領域分割と占有予測の評価です。

この研究の面白いところ

物体ごとの移動と出現・消失を時間方向に表現し、事前に付けた3次元境界枠なしで疑似ラベルを作ります。確信度と再現率の関係も評価しています。

どこまで分かった?

要旨で示された比較はSemanticKITTI上の実験です。他の環境への一般化や実運用での性能は、この要旨だけからは分かりません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

2次元画像向けの基盤モデルは3次元の意味ラベルを自動生成する手掛かりになるが、その知識を頑健な3次元表現へ移すには、通常は複雑な経験則や複数モデルの組み合わせが必要となる。本研究は、4次元ガウス表現を使い、予測の確信度を伴うLiDAR点群の領域分割結果と、任意のボクセル解像度での意味付き占有格子を取り出す自動処理系SplatLabelを提案する。 中心となるのは、個々の3次元要素の軌跡と存続期間をモデル化する明示的な時間多様体である。これにより、動く対象を追跡し、物体が現れる時点と消える時点を厳密に定め、事前に付けられた3次元境界枠を不要にする。動的な追跡を支えるため、構造と意味の事前情報も用いる。観測されない領域の形状は、360度LiDARを仮想的な深度画像として統合して補い、領域固有のプロンプト設計に頼る代わりに、2次元モデルの連続的な確率値を直接蒸留して、時間と空間にまたがる意味の曖昧さを扱う。 さらに、精度と再現率の現実的な兼ね合いを反映するため、疑似ラベル評価を選択的分類として捉え直し、一般化したリスク・再現率指標を使う。SemanticKITTIでの実験では、複数の再現率水準にわたり、既存の先進的な比較手法を一貫して上回った。3次元LiDAR領域分割と占有予測の双方に向けた頑健な枠組みとなることを示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

While 2D Vision Foundation Models offer a pathway to automate 3D semantic pseudo-labelling, translating these priors into robust 3D representations typically requires complex heuristics or multi-model ensembles. We introduce SplatLabel, an automated pipeline that leverages a 4D Gaussian representation to extract LiDAR segmentation with predictive confidence, as well as semantic occupancy grids at arbitrary voxel resolutions. At its core, SplatLabel handles dynamic environments through an explicit temporal manifold that models the trajectories and lifespans of individual 3D primitives. This allows the system to accurately track moving actors and strictly define when objects appear and disappear, completely eliminating the need for pre-annotated 3D bounding boxes. To robustly support this dynamic tracking, the representation is grounded by structural and semantic priors: we guide scene geometry in unobserved regions by integrating 360-degree LiDAR via virtual depth maps, and rather than relying on domain-specific prompt engineering, we directly distill continuous soft probabilities from 2D models to inherently resolve semantic ambiguities over time and space. Finally, to accurately reflect the real-world trade-off between precision and recall, we reframe pseudo-label evaluation as a selective classification task using a generalized risk-recall metric. Experiments on SemanticKITTI demonstrate that SplatLabel consistently outperforms state-of-the-art baselines across multiple recall levels, establishing a highly robust framework for both 3D LiDAR segmentation and occupancy prediction.

arXiv ID: 2609.29836 / 要約の誤りについて