arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

2台のカメラ間で滑らかにズームする拡散モデル

ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming

Jiayi Zhang, Renlong Wu, Yukang Ding, Sibin Deng, Wangmeng Zuo

この論文をやさしく読む

ひとことで言うと

2台のカメラを切り替えながらズームする際の形や色の飛びを、拡散モデルで抑える方法を提案した。

何に役立つ?

スマートフォンなどの複数カメラによるズーム映像の改善に役立つ可能性がある。実証は合成・実世界データでの手法比較である。

この研究の面白いところ

潜在空間と画素空間の両方を使い、幾何の整合性、高周波の細部、時間的な滑らかさをそれぞれ補う。

どこまで分かった?

要旨は合成・実世界データで従来法を上回ったと述べるが、具体的な改善数値や端末上での処理速度は記載していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

2台のカメラを切り替えるデジタルズームでは、幾何学的な形や色の一貫性に目立つ不連続が生じ、利用体験が損なわれることが多い。近年の二眼カメラ間の滑らかなズーム(DCSZ)手法は、フレーム補間モデルをDCSZデータで追加学習して改善を図るが、視点間の大きな差や複雑な幾何変換への対応が難しい。拡散モデルの生成に関する事前知識がこの問題に適すると考え、DCSZへの応用を調べる。ただし、既存の拡散モデルによるフレーム補間を単純に適用しても、条件付けの不足、VAE符号化中の高周波情報の損失、時間的一貫性の不足により、切り替え映像の忠実度は低い。そこで、潜在空間と画素空間の両方で二眼カメラ入力を活用し、写真のように自然な切り替えを作る高忠実度の拡散モデルZoomDiffを提案する。まず、多段階のノイズ除去中に2枚の画像による条件付けを強め、幾何学的一貫性を改善する。次に、VAEエンコーダーから得た、フローに沿って位置合わせした複数のスケールの特徴をVAEデコーダーに注入し、高周波の細部を復元する。このとき、フローで誘導する時間的一貫性の教師信号も導入し、より滑らかな切り替えを生成する。合成データと実世界データの両方を使った広範な実験で、ZoomDiffは最先端の手法を定量的にも定性的にも上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle with the large cross-view disparities and complex geometric transformations. Considering that the generative prior of diffusion models is suitable for addressing this problem, we explore their application to DCSZ. However, naively applying existing diffusion-based FI models still yields low-fidelity transitions due to insufficient conditional guidance, high-frequency information loss during VAE encoding, as well as inadequate temporal consistency. To address this, we propose ZoomDiff, a high-fidelity diffusion model that leverages dual-camera inputs in both latent and pixel spaces for photo-realistic transitions. Specifically, we first strengthen dual-image conditional guidance during the multi-step denoising process to improve geometric consistency. Then we inject flow-aligned multi-scale features from the VAE encoder into the VAE decoder to recover high-frequency details, where flow-guided temporal consistency supervision are introduced to produce more smooth transitions. Extensive experiments on both synthetic and real-world datasets demonstrate that ZoomDiff outperforms state-of-the-art methods quantitatively and qualitatively. Project page: https://jiayi-hit.github.io/ZoomDiff.github.io/.

arXiv ID: 2609.28083 / 要約の誤りについて