arXiv論文メモ
新着一覧
cs.LG / cs.AI / stat.CO · 査読状況未確認

特徴量の依存を考慮して固有の予測情報を測る

MCIR: A Feature Dependence-Aware Explainability Method with Reliability Guarantees

Poushali Sengupta, Sabita Maharjan, Frank Eliassen, Shashi Raj Pandey, Yan Zhang

この論文をやさしく読む

ひとことで言うと

よく似た特徴量が複数あるとき、それぞれが他の特徴量にはない予測情報をどれだけ持つかを測る方法です。

何に役立つ?

重複や強い相関のある入力を使うモデルで、特徴量の重要度を診断するために役立ちます。少量の説明用データで計算負担を減らす方法も評価しています。

この研究の面白いところ

単独の関連の強さではなく、依存する近傍を知った後にも残る情報量を使います。完全な条件付き冗長性なら母集団スコアが0になる性質があります。

どこまで分かった?

実データでは比較基準によって優劣が分かれ、一律に既存手法を上回ったわけではありません。要旨で最も明瞭な利点が示されたのは、ほぼ重複した変数を追加した条件です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

現代の機械学習モデルには、互いに強く依存する、あるいは冗長な特徴量が含まれることが多い。共有される予測情報が相関した予測変数の間に分散するため、特徴量への寄与の割り当ては難しくなる。その結果、SHAP、LIME、HSIC、MI/CMI、SAGEなどの既存手法では、多重共線性やほぼ重複した予測変数があると、順位が不安定になる可能性がある。 本研究では、依存関係を考慮する大域的な特徴量重要度の手法として、Mutual Correlation Impact Ratio Method(MCIR-M)を提案する。これは、選択した依存近傍を超えて、各特徴量が固有に提供する予測情報を定量化する。MCIR-MはMutual Correlation Impact Ratio(MCIR)を導入し、各特徴量について強く依存する近傍を条件に置き、条件付き情報量とブロック全体の情報量の比を正規化して計算する。母集団におけるスコアは[0,1]に入り、条件付きで完全に冗長な場合には0となる。また、利用可能なデータの一部を使ってMCIRを計算し、全データによる説明との一致度を評価する軽量な推定手続きを導入する。 冗長性を制御した合成データ実験とUCI HARベンチマークにおいて、MCIRは依存関係を考慮した順位付けを示し、ほぼ重複した予測変数を追加した場合に最も明確な利点が見られた。独立SHAPと条件付きSHAP、SAGE、HSIC、相互情報量に基づくスコア、CIR系列のベースラインとの比較では、実データの評価基準によって優劣が分かれた。評価した構成では、説明に使う標本数を減らすことで計算負担が下がった。一方、全データによる説明との一致度は、順位、上位集合、忠実度の診断によって別途評価した。総じて、MCIR-Mは、特徴量間の依存が強い場合の大域的説明に対して、依存関係を考慮した実用的な診断手法を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Modern machine-learning models often contain strongly dependent or redundant features, making feature attribution difficult because shared predictive information can be distributed across correlated predictors. Existing methods such as SHAP, LIME, HSIC, MI/CMI, and SAGE may therefore produce unstable rankings under multicollinearity or near-duplicate predictors. We propose the Mutual Correlation Impact Ratio Method (MCIR-M), a dependence-aware global feature-importance approach that quantifies the unique predictive information contributed by each feature beyond a selected dependence neighbourhood. MCIR-M introduces the Mutual Correlation Impact Ratio (MCIR), which conditions each feature on strongly dependent neighbours and computes a normalized ratio of conditional to block-level information. The population score lies in [0,1] and equals zero under exact conditional redundancy. We also introduce a lightweight estimation procedure that computes MCIR using a fraction of the available data and evaluates agreement with full-data explanations. Across controlled synthetic redundancy experiments and the UCI HAR benchmark, MCIR shows dependence-aware ranking behaviour, with its clearest advantage under injected near-duplicate predictors. Comparisons with independent and conditional SHAP, SAGE, HSIC, MI-based scores, and CIR-family baselines are mixed across real-data criteria. Reduced explanation samples lower computational burden in the evaluated configurations, while agreement with full-data explanations is assessed separately through ranking, head-set, and faithfulness diagnostics. Overall, MCIR-M provides a practical dependence-aware diagnostic for global explanation under strong feature dependence.

著者のコメント

Accepted for publication in Transactions on Machine Learning Research (TMLR)

arXiv ID: 2610.01641 / 要約の誤りについて