arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

3次元構造を表す潜在空間で動画生成の整合性を改善

GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation

Jiahao Lu, Minghao Yin, Wenbo Hu, Hengyu Liu, Wang Zhao, Sai-Kit Yeung, Ying Shan, Yuan Liu

この論文をやさしく読む

ひとことで言うと

動画生成の内部表現を、見た目だけでなく奥行きやカメラ位置も復元できる形に変える研究です。

何に役立つ?

考えられる用途は、視点が変わっても形が矛盾しにくいシーン生成です。2つのデータセットで画質指標などを評価しています。

この研究の面白いところ

生成器と学習手順を固定し、潜在空間の置き換えによる効果を比較しています。画質とは別に3次元整合性も測っています。

どこまで分かった?

FVDの低下率はデータセットごとに異なり、軌跡誤差半減はRealEstate10Kの結果です。あらゆるシーンで同じ改善を得ることを示したものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

知覚と生成に共通する基盤として、幾何構造を本来の表現に組み込んだコンパクトな潜在空間を提案する。視覚生成器は写実的なフレームを作れても、一貫した3次元シーンを維持できるとは限らない。これはモデルだけでなく表現の問題でもあると考える。通常、生成器は見た目を中心とする潜在表現を変化させる一方、知覚モデルは複数視点間の構造を符号化した、意味的に豊かな空間で幾何構造を復元するからである。幾何構造を追加の出力にするのではなく、幾何基盤モデルの特徴を生成用のコンパクトな潜在空間へ再パラメータ化する。この転換を実現するのが幾何構造を組み込んだオートエンコーダGAEであり、その潜在表現から見た目、深度、カメラ、点マップを一体として復号できる。この状態表現を使えば、標準的な条件付きフローで多様な生成タスクに対応できる。生成器と学習手順を固定した比較では、潜在表現をGAEに置き換えることで、視覚品質と独立に測定した3次元整合性の両方が向上した。FVDはRealEstate10Kで12.7%、DL3DVで23.1%低下し、RealEstate10Kでのカメラ軌跡誤差は半減した。これらの結果は、潜在空間が幾何的に一貫した生成の中心的要素であり、知覚と生成をつなぐ共通のインターフェースになり得ることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We present a compact geometry-native latent space as a shared foundation for perception and generation. Visual generators can produce photorealistic frames without preserving a consistent 3D scene. We argue that this is not only a modeling problem but also a representation problem: generators typically evolve appearance-centric latents, while perception models recover geometry in a semantically rich space that encodes cross-view structure. Rather than adding geometry as another output, we reparameterize a geometry foundation model's features into a compact latent space for generation. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by $12.7\%$ and $23.1\%$ on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.

著者のコメント

Project page: https://jiah-cloud.github.io/GAE.github.io/ Github: https://github.com/TencentARC/GAE-GeometricAutoEncoder

arXiv ID: 2609.24981 / 要約の誤りについて