arXiv論文メモ
新着一覧
cs.RO / cs.CV · 査読状況未確認

カメラとLiDARを組み合わせたロボットの新視点画像生成

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

Yang Zhou, Jiuhong Xiao, Shizhao Ye, Long Quang, Carlos Nieto-Granda, and Giuseppe Loianno

この論文をやさしく読む

ひとことで言うと

カメラ画像とLiDARの点群を組み合わせ、ロボットが別の位置から見た画像と深度を生成する方法。

何に役立つ?

ロボットの視点移動に伴う見え方と三次元構造を予測する研究に役立つ。地上ロボットでの実環境動作も要旨に示されている。

この研究の面白いところ

画像用と点群用の事前学習モデルをつなぐための大型翻訳モデルを作らず、カメラ投影後に共通する空間構造を利用する。

どこまで分かった?

定量的な改善はGrandTourと同じ基盤モデルの画像のみの版との比較による。要旨にはあらゆる環境での精度や走行性能は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボットの新しい視点からの画像生成では、見た目と距離を伴う三次元構造の両方を復元する必要がある。しかし、生成型の既存手法の多くは画像だけを使い、ロボットに一般的な補助センサーであるLiDARを見落としている。本研究は、カメラとLiDARのマルチモーダル表現を用いる生成型の新視点画像生成法M3GDを提案する。別々に事前学習した二次元画像モデルと三次元点群モデルを組み合わせ、両者を翻訳するモデルを別途事前学習する必要がない。 カメラへの投影後、固定したLiDAR特徴と画像特徴には、共通する空間構造が相当程度あることを示す。M3GDはこの構造を使い、明示的な幾何統計量と学習済みの点群記述子を、求める視点に合わせた情報のまとまりとして画像の潜在表現の格子に配置する。これを軽量な残差アダプター経由で複数視点のフローマッチング生成器へ入力する。生成器の潜在空間、デコーダー、学習目標はそのまま保つ。 GrandTourデータセットでは、同じ基盤モデルの画像のみを使う版より、目標視点のRGB画像と深度の生成が改善した。構成要素を除く実験は、改善が画素に位置合わせされたLiDAR情報に由来すること、目標視点のLiDARが、要求された視点と元の観測を結ぶ幾何学的な問い合わせとして働くことを示した。地上ロボットへの搭載では実環境での動作を示し、Euler積分のステップ数で品質と計算費用の兼ね合いを調整できる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independently pretrained 2D image and 3D point-cloud foundation models without separately pretraining a cross-modal translator. We show that, after camera projection, frozen LiDAR and image features exhibit substantial shared spatial structure, providing a natural cross-modal representation. M3GD conditions generation on LiDAR through this structure: it combines explicit geometry statistics with learned point-cloud descriptors into view-aligned packets on the image-latent grid, injected through a lightweight residual adapter into a multi-view flow-matching generator whose latent space, decoders, and training objective remain intact. On the GrandTour dataset, M3GD improves target-view RGB and depth synthesis over an image-only version of the same backbone. Ablations show that the gains come from pixel-aligned LiDAR content and that target-view LiDAR acts as a geometric query linking the requested view to source observations. Deployment on a ground robot demonstrates practical real-world operation, with a configurable quality--cost trade-off controlled by the number of Euler integration steps.

arXiv ID: 2609.30056 / 要約の誤りについて