情報量の最適化だけでは分布変化への頑健性は保証されない
On the Limits of Maximal Coding Rate Reduction for Out-of-Distribution Generalisation
この論文をやさしく読む
ひとことで言うと
学習した表現が情報理論的に最適でも、環境が変われば予測がほぼ全部外れる場合があると示した研究です。
何に役立つ?
解釈しやすい目的関数や環境間で共通の最適性を、安全性や汎化の保証と同一視しないための理論的な判断材料になります。
この研究の面白いところ
訓練では見られない入力が原因となる例だけでなく、テスト入力がすべて訓練時にも起こり得る状況でも、ほぼ100%の誤りを構成しています。
どこまで分かった?
これはMCR²が失敗し得ることを示す限界分析です。すべての応用で失敗するという結論や、実際の失敗頻度を測定した結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
深層学習の目的関数、表現、構造を解釈可能にするため、多くの取り組みがなされてきた。その目的は、多様な実世界の用途で学習システムの安全性、頑健性、汎化性能を向上させることである。近年提案された最大符号化率削減(MCR²)は、クラスごとの部分多様体に対し、構造化され識別力のある表現を学ぶ有望な情報理論的枠組みであり、解釈可能なホワイトボックス構造にも着想を与えてきた。しかし、MCR²は分布が変化すると完全に失敗し得ることが観察される。これを動機として、分布外(OOD)汎化の限界を調べる。 OOD汎化に関するMCR²の二つの限界を示す。第一に、MCR²の目的関数だけでは、予測の完全な失敗が起こり得る。完全に安定した特徴を利用できるにもかかわらず、不安定な環境特徴だけに基づく表現が符号化の大域的最適値を達成し、相関が反転すると完全に失敗する場合がある。この厳密な最適解の例には、訓練時には生じ得ないテスト入力が含まれる。すべての可能なテスト入力が訓練時にも生じ得る場合でも、符号化の質を最適値に任意に近づけながら、予測誤り率を100%に任意に近づけることができる。 第二に、広く成功を収めている不変リスク最小化(IRM)とリスク外挿(REx)の基礎となる不変性の原理を直接組み込んでも、この失敗は解消されない。失敗する表現は、複数の訓練環境で同一の最適な符号化作用素を持ち得る。これは、符号化の最適性を共有しても安定した予測は保証されないことを示している。したがって、MCR²について信頼できるOOD保証を得るには、環境をまたいで安定した予測関係を確立する、新たな追加仮定や学習原理が必要である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Substantial efforts have been devoted to making deep learning objectives, representations, and architectures interpretable, with the goal of improving the safety, robustness, and generalisation of learning systems in diverse real-world applications. The recently proposed maximal coding rate reduction ($\mathrm{MCR}^{2}$) offers a promising information-theoretic framework for learning structured, discriminative representations of class-wise submanifolds and has inspired interpretable white-box architectures. However, we observe that $\mathrm{MCR}^{2}$ can completely fail under distribution shift, motivating our study of its out-of-distribution (OOD) generalisation limits. We establish two limitations of $\mathrm{MCR}^{2}$ for OOD generalisation. First, the $\mathrm{MCR}^{2}$ objective alone can admit complete prediction failure: a representation based entirely on unstable environmental features can achieve the global coding optimum yet fail completely after correlation reversal, despite an available perfectly stable feature. This exact-optimum example includes test inputs that cannot occur during training. Even when every possible test input can also occur during training, coding quality can be arbitrarily close to optimal while prediction error is arbitrarily close to 100%. Second, directly incorporating the invariance principle underlying widely successful invariant risk minimisation (IRM) and risk extrapolation (REx) does not eliminate this failure. The failing representation admits the same optimal coding operator across training environments, showing that shared coding optimality does not ensure stable prediction. Reliable OOD guarantees for $\mathrm{MCR}^{2}$ therefore require additional new assumptions or learning principles that establish stable predictive relationships across environments.
著者のコメント
31 pages, 8 figures, including appendices
arXiv ID: 2609.21001 / 要約の誤りについて