arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

隠れた物体を3D生成モデルで補い場面を分解再構成

DecomVoxel: Harnessing 3D-Native Priors with Guided In-situ Denoising Optimization for Decompositional Scene Reconstruction

Junfeng Ni, Zirui Zhou, Yixin Chen, Yu Liu, Nan Jiang, Zhifei Yang, Song-Chun Zhu, and Siyuan Huang

この論文をやさしく読む

ひとことで言うと

物が重なって見えない部分を3D生成モデルで補いつつ、元の位置や形に合うよう誘導し、物体と背景を分けて再構成します。

何に役立つ?

考えられる用途は、遮蔽のある室内などから、個別に扱えるテクスチャ付き3D物体を作ることです。要旨ではReplicaとScanNet++での評価が報告されています。

この研究の面白いところ

3Dモデルが生成した形を後から置くだけでなく、元の場面の占有位置と空位置を使って、ノイズ除去の最中に位置と形を誘導しています。

どこまで分かった?

隠れた部分は生成による補完を含むため、再構成された形が実際の不可視部分そのものだと確認されたわけではありません。要旨には具体的な評価値や失敗例は記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

分解型の場面再構成は、物体と背景を高品質に再構成することを目指すが、既存手法は強い遮蔽がある場合の品質に依然として課題がある。生成モデルの事前知識は解決策となり得るものの、2次元画像に基づく事前知識は3次元の認識が欠けているため、多視点での不整合を起こしやすい。逆に、3次元を直接扱う事前知識は、構造についてより強い帰納バイアスを与えるが、複雑な場面では空間的なずれや位置合わせの不一致を引き起こすことが多い。 これらの問題に対処するため、本研究ではDecomVoxelを提案し、物体の補完を、その場での誘導付きノイズ除去最適化として定式化することで、3次元を直接扱う事前知識とニューラル場面再構成を結び付ける。この枠組みは、安定した潜在表現の洗練を確保するため、再定式化したεベースの蒸留損失を導入する。併せて、占有位置と空位置のアンカーを時間的なアニーリングとともに用いる適応的な空間誘導を導入し、生成による架空の構造を抑え、空間的なずれを軽減する。 ReplicaとScanNet++での実験は、DecomVoxelが元の空間配置、構造の忠実性、様式の整合したテクスチャを保ちながら、最先端の手法を大きく上回ることを示す。本手法は、整ったトポロジー、幾何形状、外観を備えた高品質なテクスチャ付きメッシュを提供することで、分解型再構成の到達範囲を広げ、複雑な実世界の場面を分解して再構成するための頑健な解決策を提供する。コードはhttps://github.com/DecomVoxel/DecomVoxelで公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Decompositional scene reconstruction aims to reconstruct high-quality objects and background, yet existing methods still struggle with the level of quality under heavy occlusions. While generative priors offer a potential solution, 2D image-based priors often suffer from multi-view inconsistency due to a lack of 3D awareness. Conversely, 3D-native priors provide stronger structural inductive biases but frequently lead to spatial drift and misalignment within complex scenes. To address these issues, we propose DecomVoxel, formulating object completion as a guided in-situ denoising optimization that bridges 3D-native priors with neural scene reconstruction. Our framework introduces a reformulated epsilon-based distillation loss to ensure stable latent refinement, alongside adaptive spatial guidance that utilizes occupied and vacant anchors with temporal annealing to suppress generative hallucinations and mitigate spatial drift. Experiments on Replica and ScanNet++ show that DecomVoxel significantly outperforms state-of-the-art methods while faithfully preserving the original spatial layout, structural fidelity, and style-consistent texture. Our method pushes the boundary of decompositional reconstruction by delivering high-quality textured meshes with clean topology, geometry, and appearance, providing a robust solution for the decompositional reconstruction of complex real-world scenes. Code is available at https://github.com/DecomVoxel/DecomVoxel.

著者のコメント

SIGGRAPH Asia 2026 - Journal Track (TOG). Project page: https://decomvoxel.github.io/DecomVoxel-Webpage/

arXiv ID: 2610.01914 / 要約の誤りについて