arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

鏡像の位置と見た目を合わせる画像生成法

PhysReflect: Geometry and Perception Guided Diffusion for Physically-Plausible Mirror Reflections

Shuheng Ge, Hongwei Ren, Li Zhang, Xiangqian Wu

この論文をやさしく読む

ひとことで言うと

画像生成で鏡の中の物体が不自然にずれたり変形したりする問題に対し、幾何と見た目の両方を学習時に確認する。

何に役立つ?

鏡を含む画像の生成品質を改善する用途が考えられる。

この研究の面白いところ

鏡像の対応と境界に加え、物体の同一性や照明まで別の損失で扱い、合成画像と実画像で評価した。

どこまで分かった?

要旨はベンチマークでの優位を述べるが、数値やすべての種類の鏡面で成立する条件は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

拡散モデルは高品質な画像を作れる一方、鏡面反射の物理的な規則を破り、位置のずれ、向きの不一致、比率の不均衡、形のゆがみが生じる。既存法は合成データの増加や深度の条件付けで対応するが、潜在空間のノイズ再構成損失だけによる間接的な監督では、反射固有の幾何学的・知覚的制約を直接課せない。提案するPhysReflectは、学習の各段階で予測したクリーンな潜在表現を画素空間へ復号し、二つの微分可能な目的関数を段階的に適用する。幾何損失は、疎なエピポーラ対応と密な境界投影の整合により鏡像の空間的一貫性を促し、SAM2に基づくTwinTrackが鏡の中の位置を安定して捉える。知覚損失は、DINOv2の特徴による意味的一貫性で反射された物体の同一性と外観を保ち、単眼の幾何学的事前知識を使う照明的一貫性で深度、面法線、照明の整合を促す。合成および実世界のベンチマーク実験では、幾何、知覚、物理的なもっともらしさの指標と定性的な見た目で、従来の鏡面反射生成法を上回ったと報告する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Diffusion models generate high-quality images, yet often violate the physical laws governing mirror reflections. Reflections often suffer from geometric aberrations, including positional offsets, directional misalignment, proportional imbalance, and structural distortion. These failures remain evident even in contemporary state-of-the-art generative systems. Existing methods itigate this problem through synthetic data scaling or auxiliary depth conditioning, yet their merely reliance on latent-space noise reconstruction losses as implicit supervision prevents direct enforcement of reflection-specific geometric and perceptual constraints. To bridge this gap, we present PhysReflect, a geometry and perception guided diffusion framework that decodes the predicted clean latent into pixel space at each training step and applies annealed supervision through two complementary differentiable objectives. The Geometric Loss enforces mirror-induced spatial consistency through sparse epipolar correspondence and dense boundary projection alignment, where a SAM2-based TwinTrack mechanism provides stable in-mirror localization for boundary-aware supervision. The Perceptual Loss preserves reflected appearance by combining Semantic Consistency Loss, which maintains reflected identity and appearance via DINOv2 features, and Lighting Consistency Loss, which regularizes depth, surface-normal, and illumination coherence under monocular geometry priors. Experiments on synthetic and real-world benchmarks show that PhysReflect outperforms prior mirror-reflection methods in geometric, perceptual, and physical-plausibility metrics, as well as qualitative visual results.

arXiv ID: 2609.23442 / 要約の誤りについて