arXiv論文メモ
新着一覧
stat.ME · 査読状況未確認

一般化線形モデルの変数選択に使う滑らかな情報量規準

Smooth Information Criterion for Variable Selection in Generalised Linear Models

Andrew McInerney

この論文をやさしく読む

ひとことで言うと

変数の組み合わせを総当たりせず、BICが選ぶモデルに近い結果を滑らかな最適化で求める方法です。

何に役立つ?

一般化線形モデルで多くの候補変数から説明に必要なものを選ぶ際、計算負担を抑える選択肢になります。

この研究の面白いところ

ガウス、二項、ポアソン回帰で全探索のBIC選択に近く、16変数の実データでは65,536通りの中の最適モデルを見つけています。

どこまで分かった?

要旨での全探索との直接比較は計算可能な場合に限られます。予測性能は他手法より一貫して高いとは述べられておらず、おおむね同程度です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

情報量規準を使った変数選択には明確な統計的な目標があるが、候補モデルの離散的な探索が必要になる。滑らかな情報量規準(SIC)は、不連続なモデル次元の項を微分可能な近似で置き換え、εを段階的に変える継続法によってこの近似を徐々に鋭くする。正則化の強さをデータから選ぶ必要はない。本研究は、一般化線形モデル(GLM)で係数単位の変数選択を行う一般的な手法としてSICを発展させ、滑らかな最適化が対応する離散的な情報量規準の選択問題をどれだけ忠実に再現するかを調べる。BICに焦点を当て、可能な場合は全変数部分集合の探索と直接比較する。評価には、選択した変数集合の完全一致、BICの差、信号の強さを変えた境界付近での選択挙動を用いる。ガウス回帰、二項回帰、ポアソン回帰のシミュレーションでは、SICは全探索によるBIC選択をよく再現し、正確なBIC選択境界にも追随した。逐次選択によるBIC、LASSO、SCAD、MCPとの比較では、疎なモデルを維持しながら競争力のある変数選択性能を示し、予測性能は各手法でおおむね同程度だった。説明変数の次元が増えるほど、逐次BICに対する計算上の利点が大きくなった。候補の説明変数が16個ある実データでは、65,536通りすべての変数集合の中でBICが全体最適となるモデルを、全探索よりはるかに少ない計算量で見つけた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Variable selection using information criteria has an explicit statistical target but requires discrete search over candidate models. The smooth information criterion (SIC) replaces the discontinuous model-dimension term by a differentiable approximation, with an $\epsilon$-telescoping continuation strategy progressively sharpening this approximation without data-driven selection of a regularisation-strength parameter. We develop SIC as a general procedure for coefficient-level variable selection in generalised linear models (GLMs) and use it to address a central question: how faithfully does smooth optimisation reproduce the corresponding discrete information-criterion selection problem? Focusing on BIC, we benchmark SIC directly against exhaustive subset selection where feasible, using exact support agreement, BIC difference and selection behaviour across a varying signal-strength boundary. Simulations in Gaussian, binomial and Poisson regression show that SIC closely reproduces exhaustive BIC selection and tracks the exact BIC selection boundary. Relative to stepwise BIC, LASSO, SCAD and MCP, SIC produces competitive variable-selection performance while retaining sparse models, with predictive performance broadly comparable across methods. Computational advantages over stepwise BIC increase with predictor dimension. In a real-data application with 16 candidate predictors, SIC recovers the globally BIC-optimal model among all 65,536 supports at a small fraction of the computational cost of exhaustive enumeration.

arXiv ID: 2609.30174 / 要約の誤りについて