arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

3D生成物の色と質感を別々の参照画像から変更する

DiDE:Direct Injection with Color-Texture DEcoupling for 3D Stylization

Tao Wu, Alexandra Gomez-Villa, Senmao Li, Yaxing Wang, Joost van de Weijer, Kai Wang

この論文をやさしく読む

ひとことで言うと

3D物体の形を保ちながら、色と質感を別々の画像から指定する方法です。既存の画像から3Dを作るモデルを使い、追加学習なしで二つの属性を分けて扱います。

何に役立つ?

考えられる用途は、3Dアセット制作で色だけ、質感だけ、または両方を別々に調整することです。複数参照画像のベンチマークで転写と形状保持を評価しています。

この研究の面白いところ

モデル内部の質感情報が使うチャネルが一部に限られる点を利用し、残りの空間へ色を独立に入れます。各自己注意層で信号を合成する設計です。

どこまで分かった?

実験結果の具体的な改善値やベンチマーク規模は要旨にありません。追加学習不要という主張は既存の事前学習モデルを利用する設定であり、学習済みモデル自体が不要という意味ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

整流フローに基づく画像から3Dを生成するモデルの近年の進展により、高忠実度の3Dアセット生成が可能になった。これを基盤として、強力な3Dの事前知識を利用し、参照画像の視覚属性を生成した3Dアセットへ移す、追加学習を要しないスタイル変換の研究が増えている。しかし既存手法では、色と質感を一緒に移すか、どちらも移さないかの方式に限られ、両者を独立に制御する仕組みがない。本研究では、この制約に対する課題を、分離型3Dスタイル変換(Disen3D)として定式化する。 この課題に対処するため、Disen3Dのための初の追加学習不要な枠組みDiDEを提案する。中心となるのは、画像から3Dを生成するモデルの構造化された潜在空間が、質感に対して過完備であるという観察である。質感情報が占めるのは、スタイルに重要なチャネルのごく一部だけであり、独立した色の符号化に使える自由な部分空間が残っている。DiDEはこれをチャネル分割の仕組みで活用する。内容を表す画像、質感の参照画像、色の参照画像を専用の分岐で処理し、内容の形状を全体を通して保ちながら、すべての自己注意層で二つのスタイル信号を干渉させずに合成する。新たに収集した複数参照のベンチマークDisen3D-Benchでの実験では、色の忠実さ、質感の転写、内容の保持において、DiDEが2Dおよび3Dのスタイル変換ベースラインを一貫して上回る。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently -- a limitation we formalize as Disentangled 3D Stylization(Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.

著者のコメント

Accepted to NeurIPS 2026

arXiv ID: 2610.02044 / 要約の誤りについて