arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

単一画像から編集・シミュレーション可能な3D場面を生成

Fysiverse-3D-Vision Technical Report: Generating Executable 3D Worlds from Images through Unified Spatial Reasoning

Dingkang Yang, Yizhou Liu, Wendong Cheng, Zizhi Chen, Shunli Wang, Yang Liu, Hongsheng Li, Lihua Zhang

この論文をやさしく読む

ひとことで言うと

一枚の画像から物体の形だけでなく位置・向き・大きさを推定し、編集やシミュレーションに使える3D場面を作る研究。

何に役立つ?

対話的な3D編集や物理シミュレーション向けに、画像から場面全体を再構成する方法の参考になる。

この研究の面白いところ

物体の生成と場面内の配置推論を切り分け、言語・視覚・幾何の情報を一つのモデルで扱う。

どこまで分かった?

要旨は比較実験での優位を述べるが、各指標の数値や実際の物理シミュレーションでの成功率は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

画像を条件とする三次元コンテンツ生成は進歩しているが、一枚の画像から制御でき、実行可能な三次元場面を作ることは依然難しい。既存の三次元生成法は見た目がもっともらしい物体や場面を合成できる一方、空間配置の推定が特定の物体生成器に結び付いている。対話的な編集、物理シミュレーション、身体性を持つ応用に必要な、物体の意味、実寸に基づく幾何、場面全体の空間関係を同時にモデル化することが苦手である。本研究は、一枚の画像から生成的な三次元場面再構成と実行可能な物体データの構築を行う、視覚・言語・幾何を統一したFysiverse-3D-Visionを提案する。空間推論と幾何再構成が互いを高め合う共通の表現を作り、個々の物体生成器の制約を超えて配置を推定できるようにする。統一されたTransformerに、文章による教師信号、意味を表す視覚的手掛かり、幾何表現を統合し、場面の文脈、実寸の幾何、物体間の相互作用を捉える。物体を条件とする配置部では、対象物体の表現と場面全体の幾何特徴の間でクロスアテンションを行い、物体の平行移動、回転、尺度を予測する。学習では、まず幾何と言語の対応を学び、再構成能力を保ちながら配置の推論を加え、衝突を考慮する最適化で物理的な整合性を改善する。空間配置の推論と物体合成を分けることで、対話的な場面編集、物体単位の操作、実行可能な三次元コンテンツ生成に適応できるインターフェースを提供する。実験では、従来法に比べ、幾何の整合性、配置推定、描画品質、物理的性質の理解で優れた性能を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative models have advanced image-conditioned 3D content creation, yet generating controllable and executable 3D scenes from a single image remains challenging. Existing 3D generative approaches can synthesize visually plausible objects and scenes, but their spatial layout estimation is coupled with specific asset generators. They struggle to jointly model object semantics, metric geometry, and scene-level spatial relationships, which are essential for interactive editing, physical simulation, and embodied applications. We propose Fysiverse-3D-Vision, a unified vision-language-geometry framework for generative 3D scene reconstruction and executable asset construction from a single image. We establish a shared representation where spatial reasoning and geometric reconstruction mutually enhance each other, allowing object layouts to be inferred beyond the constraints of individual asset generators. Our model integrates textual supervision, semantic visual cues, and geometric representations within a unified Transformer to capture scene context, metric geometry, and object-level interactions. An object-conditioned layout module performs cross-attention between target object representations and global geometric features to predict object translation, rotation, and scale. Training progressively learns geometry-language alignment, introduces layout reasoning while preserving reconstruction capability, and refines physical consistency through collision-aware optimization. By separating spatial layout reasoning from asset synthesis, Fysiverse-3D-Vision provides an adaptable interface for interactive scene editing, object-level manipulations, and executable 3D content generation. Experiments demonstrate that our framework achieves superior geometric consistency, layout estimation, rendering quality, and physical property understanding compared with existing approaches.

著者のコメント

Fysics AI Technical Report

arXiv ID: 2609.25741 / 要約の誤りについて