arXiv論文メモ
新着一覧
eess.IV · 査読状況未確認

画像品質指標を動画符号化のブロック最適化に使う方法

Rate-distortion optimization for full-reference image quality metrics via stochastic Hessian estimates

Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega

この論文をやさしく読む

ひとことで言うと

人の見た目により近い画像品質指標を、動画符号化器のブロックごとの設定選択に使えるよう近似した。

何に役立つ?

動画の符号量と画質の兼ね合いを、SSE以外の品質指標で最適化したい場合の方法になる。復号器を変えずに評価した結果が示されている。

この研究の面白いところ

画像全体が必要な指標をヘッセ行列の二次形式で近似し、さらにブロック対角または対角成分だけにして符号化処理に組み込んだ。

どこまで分かった?

示された改善はKodakとCLIC、VVC、五つの指標での評価に基づく。符号化の計算量は10~30%増えた。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ブロック単位の動画符号化器は、符号量と画質劣化の兼ね合いを最適化して入力に応じた符号化パラメーターを選ぶ。従来の劣化指標である二乗誤差和(SSE)は、各ブロックのSSEの和になるため、符号量と劣化の最適化をブロックごとに独立して行いやすい。一方、MS-SSIMやLPIPSなど、元画像との比較に基づく画像品質評価指標は、SSEより人間の視覚に合う場合が多いが、ブロック単位に分解できず、通常は復号した画像全体を入力とするため、符号化の処理中に直接使えない。 本研究は、指標の二次近似に関する既存の結果を基に、広い種類の画像品質評価指標を、入力画像に依存する二次形式の劣化指標(IDQD)で近似する。その二次形式の行列は、元動画について評価した品質指標のヘッセ行列から導く。ブロック単位で劣化を計算できるよう、ヘッセ行列のブロック対角部分だけを残す方法と、対角部分だけを残す方法を提案する。どちらも自動微分で得たヘッセ行列とベクトルの積だけで計算できる推定法を示す。 VVCでKodakとCLICのデータを用い、五つの品質指標について評価したところ、IDQDを用いた符号量・劣化最適化は、対象の品質指標で測ったBD-rateを14.2~36.7%削減した。復号器の変更は不要で、符号化の計算量の増加は10~30%だった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Block-based video codecs select coding parameters based on the input by optimizing a rate-distortion trade-off. The conventional distortion choice, the sum of squared errors (SSE), simplifies parameter selection: the SSE is the sum of block-wise SSEs, so rate-distortion optimization (RDO) can treat blocks independently. Alternatively, full-reference image quality assessment (FR-IQA) metrics such as MS-SSIM or LPIPS often align better with the human visual system than SSE, but they cannot be used in-loop: they do not decompose block-wise and typically require the fully decoded image as input. Building on existing results in metric quadratization, we approximate a broad class of FR-IQA metrics by an input-dependent quadratic distortion (IDQD), whose quadratic form matrix is derived from the Hessian of the metric evaluated at the source video. To make the distortion computable block-wise, we propose two approximations of the Hessian matrix: 1) keeping the block-diagonal, and 2) keeping only its diagonal. We propose estimators for both that require only matrix-vector products with the Hessian obtained by automatic differentiation. Across five metrics for Kodak and CLIC in VVC, IDQD-RDO achieves 14.2-36.7 % BD-rate savings under the target metric with no decoder changes and incurs 10-30 % encoding complexity overhead.

arXiv ID: 2609.30077 / 要約の誤りについて