3D画像と言語の処理で点群・ボクセル・複数視点を統合
From Alignment to Fusion in 3D Vision-Language
この論文をやさしく読む
ひとことで言うと
点群、ボクセル、複数視点画像の特徴を、形の関係を壊さないように揃えて統合する3D視覚・言語モデル。
何に役立つ?
3D空間内の物体分割、言葉で指定された対象の特定、質問応答、説明文生成のモデル設計に役立つ可能性がある。要旨では8データセットでの性能比較を報告している。
この研究の面白いところ
融合前に特殊直交群に制約した写像を使い、各表現内の距離と内積を保ちながら特徴を変換する。対象特定では比較法より複数データセットで正解率が向上した。
どこまで分かった?
示された改善は記載された8データセットでの実験結果である。未知の場面全般への性能は要旨からは判断できない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
統一的な3D視覚・言語システムには、インスタンス分割から言語に基づく推論まで対応しながら、幾何、スケール、外観の相補的な手掛かりを組み合わせる必要がある。従来法では点群、ボクセル格子、複数視点画像を別々に処理することが多い。異なる表現を直接合わせるだけでは特徴の差が大きく残る可能性があり、後から制約なしに適応させると各表現の内部の幾何をゆがめるおそれがある。 本研究は、まず三つの表現の各組み合わせについてコサイン類似度による整列を行い、セグメント単位の対応を作り、その後プロンプトで誘導するクエリデコーダで課題に応じた特徴を取り出す「整列してから融合する」枠組みを提案する。融合前には、表現ごとのクエリ特徴を、特殊直交群に制約された学習可能な写像で変換する。この写像は各表現内の内積とユークリッド距離を保つため、内部の幾何を恣意的にゆがめずに、表現に応じた制御可能な再パラメータ化ができる。変換した特徴を下流課題の教師信号の下でAdaptive Fusionにより結合する。 実験では、インスタンス分割、視覚的な対象特定、質問応答、密な説明文生成について、8つのデータセットを扱った。PQ3Dと比べ、ScanNet200では平均適合率が3.2ポイント、ScanRefer、Nr3D、Sr3D、Multi3DReferでは対象特定の正解率がそれぞれ2.9、10.6、4.6、4.1ポイント向上した。ScanQA、SQA3D、Scan2Capでも性能が向上した。要素を除く実験は、整列と直交再パラメータ化の相補的な役割、およびAdaptive Fusionの有効性をさらに裏付けた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous representations may leave substantial feature discrepancy unresolved, while subsequent unconstrained adaptation may distort their internal geometry. We propose an align-then-fuse framework that first applies triple pairwise cosine alignment to establish segment-level correspondence across the three representations and then retrieves task-conditioned features with a prompt-guided query decoder. Before fusion, representation-specific query features are transformed by learnable mappings constrained to the special orthogonal group. These mappings preserve inner products and Euclidean distances within each representation, permitting controlled representation-specific re-parameterisation without arbitrarily distorting its internal geometry. The transformed features are subsequently combined through Adaptive Fusion under downstream task supervision. Experiments cover eight datasets for instance segmentation, visual grounding, question answering, and dense captioning. Compared with PQ3D, the model improves average precision by 3.2 points on ScanNet200 and grounding accuracy by 2.9, 10.6, 4.6, and 4.1 points on ScanRefer, Nr3D, Sr3D, and Multi3DRefer, respectively, while also improving performance on ScanQA, SQA3D, and Scan2Cap. Ablations further support the complementary roles of alignment and orthogonal re-parameterisation and the effectiveness of Adaptive Fusion.
arXiv ID: 2609.28222 / 要約の誤りについて