最適輸送で画像の構造を保つ対照学習の正例を作る
Positive Pair Geometry Matters: Optimal Transport for Contrastive Learning of Visual Representations
この論文をやさしく読む
ひとことで言うと
同じ画像の仲間として学習させる正例を、無作為な変形だけで作らず、元画像と変形後の画像の間を最適輸送で補間して作ります。
何に役立つ?
画像の自己教師あり学習で、元の構造と整合した学習用の対を作る方法として役立ちます。エンコーダー構造を変えずに利用できるとしています。
この研究の面白いところ
画像の拡張の強さだけでなく、画素の空間分布がどのように移り変わるかを正例の設計に取り入れています。
どこまで分かった?
要旨には改善量、計算費用、個別のデータセット名はありません。構造の保持や転移性能はベンチマーク比較に基づくもので、すべての拡張で意味が完全に保たれる保証は述べられていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
対照的な自己教師あり学習は、同じ画像から拡張した複数のビューを使って表現を学ぶことで、高い性能を達成してきた。しかし、多くの既存手法は、独立にサンプリングした確率的な拡張で正例対を構成するため、意味内容が変わったり、データ分布が本来持つ幾何学を無視したりする可能性がある。本研究では、幾何学的に整合した正例を生成する、最適輸送を考慮した対照表現学習の枠組み OTCLR を提案する。 ランダムに拡張した2つのビューを直接対比する代わりに、エントロピー正則化を用いた最適輸送の変位補間によって、元画像とその拡張画像の間に中間ビューを構成する。輸送によって補間したサンプルを正例ビューとして使うことで、空間的な分布の幾何学を明示的にモデル化しながら、画像構造をよりよく保つ。さらに滑らかな表現学習を促すため、輸送補間したビューが両端の画像と整合するよう促す、補助的な Sinkhorn 正則化項を評価する。提案手法は、エンコーダーの構造を変更せず、標準的な対照学習の処理系に組み込める。複数のベンチマークデータセットでの実験は、従来の拡張に基づく対照学習の基準手法に比べ、表現の質と転移学習の性能が改善することを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Contrastive self-supervised learning has achieved strong performance by learning representations from multiple augmented views of the same image. However, most existing methods construct positive pairs using independently sampled stochastic augmentations, which may alter semantic content and ignore the intrinsic geometry of the data distribution. In this work, we propose OTCLR, an optimal transport-aware framework for contrastive learning representations that generates geometry-consistent positive samples. Instead of directly contrasting two randomly augmented views, we construct intermediate views between the original image and its augmented variants through entropic optimal-transport displacement interpolation. These transport-interpolated samples serve as positive views that better preserve image structure while explicitly modeling spatial distributional geometry. To further promote smooth representation learning, we evaluate auxiliary Sinkhorn regularization terms that encourage transport-interpolated views to remain consistent with their endpoint images. The proposed method can be incorporated into standard contrastive learning pipelines without modifying the encoder architecture. Experiments on multiple benchmark datasets show that our approach improves representation quality and transfer learning performance compared with conventional augmentation-based contrastive learning baselines.
著者のコメント
12 pages, 5 figures
arXiv ID: 2609.24125 / 要約の誤りについて