arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

人が確認できる画質を保ちながら機械向け画像を圧縮

You've Seen Enough: Quality-Constrained Image Coding for Machines

Khoa Pham-Dinh, Sanaz Nami, Hamed Rezazadegan Tavakoli, Moncef Gabbouj, Farhad Pakdaman

この論文をやさしく読む

ひとことで言うと

人が確認できる画質を目標に保ち、機械の画像認識向けに圧縮率を改善した。

何に役立つ?

考えられる用途は、人が結果を確認する画像認識システムで、転送・保存する画像の量を減らすこと。

この研究の面白いところ

画質を必要以上に上げる代わりに、容量を機械の課題性能へ振り向ける制約を置いた点。

どこまで分かった?

数値は指定された圧縮と画像分割の実験での結果であり、ほかの視覚課題で同じビットレート削減が得られるかは要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

画像は人よりもコンピュータービジョンのシステムに使われる場面が増えている。機械向け画像符号化(ICM)は、主な観察者をコンピュータービジョンのアプリケーションと考え、人はその判断を確認・検証するために画像を見ると想定する。本研究は、人にとってぎりぎり許容できる品質を定める「ちょうど気付く歪み」の考え方を参考に、人が見る品質を目標水準に抑え、残りの符号化容量を機械の性能改善に使うことを目指す。圧縮器が事前に決めた許容可能な視覚品質を満たし、残りの容量を課題用の項に配分するという制約付き最適化として、圧縮と画像分割の同時学習を捉え直す。 品質を目標へ導く罰則関数として、絶対値型と双線形型の二つを提案する。後者は目標の視覚品質を超えた場合に傾きを急にする。実験では、品質制約の下で、提案法のBD-rateが制約のないレート・歪み・課題の同時最適化に対して−22.82%、単純なレート・歪みの基準法に対して−29.81%となった。これは同じ課題性能で必要なビットレートが減ったことを示す。同時に、目標の視覚品質を妥当な誤差内で満たし、計算の複雑さも増やさなかった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Visual data is increasingly consumed by machine-vision systems rather than by human observers. Image Coding for Machines (ICM) compresses images assuming the main observer is a computer vision application and that the human observer needs to inspect or validate the decisions. Inspired by just-noticeable distortion, which sets the quality to the just-acceptable level for human observers, we aim to cap the human-observed quality at a desired level, with the goal of using the remaining coding capacity to improve the machine performance. We recast joint compression-segmentation training as a constrained optimization problem in which the codec must meet a predefined acceptable target visual quality while a task term consumes the remaining coding capacity. We solve this by designing a penalty function to guide the quality to the desired target. We propose two penalty functions, an absolute function and a bilinear function, the latter applying a steeper slope once the target visual quality is exceeded. Experimental results show that, under the quality constraint, the proposed method achieves a BD-rate of $-22.82\%$ over an unconstrained joint rate--distortion--task optimization and $-29.81\%$ over a simple rate--distortion baseline, showcasing bitrate reduction with the same task performance. This is achieved while the codec also meets the target visual quality with a reasonable error and without adding any complexity overhead.

arXiv ID: 2609.25108 / 要約の誤りについて