arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

少数の距離測定を使い単眼画像の奥行きを高速に補正

FounRef: Robust, Structure-Preserving, and Fast Metric Refinement of Frozen Monocular Foundation Priors with Sparse Anchors

Dan Halperin, Mirko Mählisch

この論文をやさしく読む

ひとことで言うと

単眼画像から得た細かな形状に、LiDARなどの少数の距離測定を合わせて、距離付きの奥行き画像を作る手法。

何に役立つ?

カメラを用いる三次元計測で、形状を保ちながら距離の尺度を与えたい場合に役立つと考えられる。基盤モデルや実測点の取得元を交換できる設計も示されている。

この研究の面白いところ

実測点をそのまま使わず、単眼モデルの密な予測と照らしてセンサー間のずれを検出する。比較実験ではDMD3Cに対し、誤差と表面ノイズを減らし、推論も大幅に速めた。

どこまで分かった?

要旨にある定量比較は、学習時と異なる領域のデータでのDMD3Cとの比較である。すべてのカメラや場面で同じ改善幅を保証する結果ではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

カメラ画像から距離の単位を持つ密な奥行きを得ることは、現実世界の三次元応用に欠かせない。しかし、精度、表面形状の忠実さ、推論速度を同時に満たすのは難しい。単眼画像の基盤モデルは転用しやすい豊かな幾何学的事前知識を持つが、信頼できる距離スケールを欠く。一方、深度補完ネットワークは距離を復元できても、形状の忠実さ、異なるデータ領域への頑健性、速度のいずれかに課題がある。 本研究のFounRefは、学習済み単眼基盤モデルを固定したまま、その奥行きの事前推定を少数の実測距離点に合わせ、密な距離付き奥行きを生成する追加学習不要の手法である。奥行きの事前モデル、実測点の取得元、補正計算器はそれぞれ独立に交換できる。実装例ではMoGe-2とLiDARの実測点を使う。まず、事前モデルの密な奥行き予測に照らして各実測点を検証し、センサー間の位置ずれに由来する、幾何学だけのフィルターでは検出できない不整合を除く。次に、細かな形状を保つ計算器で全体と局所の距離補正を行い、事前モデルが捉えた詳細な形状を残す。 FounRefは課題ごとの学習を必要とせず、未知のカメラや場面にもそのまま適用できる。学習時と異なる領域のデータでは、最新の深度補完ネットワークDMD3Cと比べ、奥行き誤差を最大24%、表面法線のノイズを92%減らし、推論速度は約15倍だった。距離スケールの調整と形状予測を切り離すことで、今後の基盤モデルや距離センサーの進歩も取り込みやすい、正確で形状に忠実かつ効率的な密な奥行き推定法を示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Dense metric depth from cameras is essential to real-world 3D applications, yet achieving accuracy, faithful surface geometry, and fast inference simultaneously remains challenging. Monocular foundation models provide rich, transferable geometric priors but lack reliable metric scale, while depth-completion networks recover metric depth at the cost of geometric fidelity, cross-domain robustness, or speed. We present FounRef, a training-free method that aligns a frozen monocular foundation prior with sparse metric anchors to produce dense metric depth. FounRef is modular by design: its depth prior, anchor source, and refinement solver can each be replaced independently. We instantiate FounRef with MoGe-2 and LiDAR anchors. FounRef validates each anchor against the prior's dense depth prediction, rejecting inconsistencies caused by cross-sensor misalignment that geometry-only filters cannot detect. It then applies global and local metric corrections through a structure-preserving solver, retaining the prior's fine-grained geometry. FounRef requires no task-specific training and operates out of the box across unfamiliar cameras and scenes. On out-of-domain data, it delivers up to 24% lower depth error, 92% lower surface-normal noise, and almost 15x faster inference than DMD3C, a state-of-the-art depth-completion network. By decoupling metric alignment from geometry prediction, FounRef provides an accurate, geometrically faithful, and efficient approach to dense metric depth that can directly benefit from future advances in foundation models and metric sensors.

著者のコメント

16 pages, 12 figures; includes appendix

arXiv ID: 2609.29224 / 要約の誤りについて