arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

ぶれた1枚の画像から任意の時刻の鮮明な画像を復元する

PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video

Junseong Shin, Hyeonsu Jo, Daehyun Kim, Tae Hyun Kim

この論文をやさしく読む

ひとことで言うと

動きでぶれた写真を、露光中の時間を平均した画像として扱い、同じモデルに時刻を指定して鮮明な1枚や動画を復元する方法です。

何に役立つ?

ぶれ除去と動画生成を別々に学習せず、1つのモデルで使い分ける設計に役立ちます。露光中央に限らず、任意の瞬間を問い合わせられる点が用途を広げます。

この研究の面白いところ

鮮明な画像だけを直接予測するのではなく、任意の時間区間の平均像を学びます。区間を組み合わせたときの整合性と、区間長ゼロの極限を同じ学習に組み込んでいます。

どこまで分かった?

最高性能という比較の範囲は、GoPro・HIDEでは生成モデル型手法、GoPro-7ではフレームごとの忠実度です。RealBlurについては競争力があるという報告です。任意のぶれで真の動画が一意に復元できるという保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動きぼけは、連続した鮮明な信号を有限の露光時間にわたって時間積分することで生じる。しかし、既存の学習手法はこの物理モデルを避け、鮮明な信号自体だけを予測している。多くの単一画像のぶれ除去手法は露光中央時刻の1フレームを復元し、ぶれ画像から動画を作る手法は固定されたフレーム集合を予測する。 本研究は、1つの決定論的モデルで、露光時間内の任意の部分区間における区間平均のぶれを直接学習する、連続時間の再定式化PickMomentを導入する。MeanFlowの平均速度の定式化になぞらえ、ぶれの積分から導く3種類の教師信号で学習する。利用可能なサブフレームからの経験的な再構成損失、重なる部分区間の間の自己整合性を課す加法性損失、区間長がゼロとなる極限を基準にした鮮明フレーム損失である。 1つの学習済みモデルは、同じネットワークへの異なる問い合わせとして、単一画像のぶれ除去、ぶれ画像からの動画生成、連続時間の好きな瞬間の復元を統合し、課題ごとの別学習を必要としない。PickMomentは、GoProとHIDEで生成モデル型のぶれ除去手法の中で最高性能を達成し、RealBlurでは復元型手法と競争力のある性能を示す。また、GoPro-7のぶれ画像からの動画生成では、フレームごとの忠実度が最も高い。いずれも反復サンプリングを使わず、1回の順伝播で実現する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video methods predict a fixed set of frames. We introduce PickMoment, a continuous-time reformulation that directly learns the interval-mean blur over arbitrary sub-intervals of the exposure with a single deterministic model. Drawing an analogy to MeanFlow's average-velocity formulation, we train the model with three supervisions derived from the blur integral: an empirical reconstruction loss from available subframes, an additivity loss that enforces self-consistency across overlapping sub-intervals, and a sharp-frame loss anchored at the zero-interval limit. A single trained model unifies single-image deblurring, blur-to-video generation, and continuous-time pick-a-moment recovery as different queries to the same network, with no separate training for each task. Our PickMoment achieves state-of-the-art performance among generative-based deblurring methods on GoPro and HIDE while competitive against restoration-based methods on RealBlur, and the highest per-frame fidelity on GoPro-7 blur-to-video, all in a single forward pass without iterative sampling.

arXiv ID: 2610.01279 / 要約の誤りについて