arXiv論文メモ
新着一覧
stat.ME · 査読状況未確認

非線形に変換した説明変数の欠測を扱う確率的EM法

A Stochastic EM Algorithm with Sampling-Importance Resampling for Missing Data in Regression with Nonlinear Predictors

Dale S. Kim

この論文をやさしく読む

ひとことで言うと

回帰分析で説明変数が欠け、さらに非線形変換を使う場合に、欠測値のサンプリングを組み込んで推定する方法です。

何に役立つ?

スプラインなどの柔軟な説明変数を使う欠測データ分析で、偏りと信頼区間の被覆を改善する用途が考えられます。

この研究の面白いところ

欠測値の条件付き分布を閉じた式で求める代わりに、完全データ尤度を比例定数まで計算できれば使える構成です。

どこまで分かった?

要旨の性能比較は2つのシミュレーションです。限界を議論するとありますが内容は要旨に示されず、欠測機構の具体的な仮定もこの要旨だけでは確認できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

説明変数に非線形変換を用いる回帰モデルでは、データが欠測していると推定が難しい。非線形性によって、欠測値の条件付き分布が通常は扱いにくくなるためである。従来の方法は、多項式や交互作用などの特定の非線形形式を必要とするか、偏りを生み得る近似に依存する。 本研究では、サンプリング・重要度リサンプリングを用いる確率的EMアルゴリズムSIR-StEMを提案する。これは、実質的な理由に基づく変換でも、スプライン基底展開のように付随して導入される変換でも、説明変数の任意の非線形変換のもとで欠測データを扱う。線形性や閉形式の条件付き分布を求める方法と異なり、SIR-StEMに必要なのは、完全データの尤度を比例定数を除いて評価できることだけであり、幅広い非線形回帰モデルに適用できる。計算効率を高めるために欠測パターンの情報を使うアルゴリズムを構成し、推定量の漸近正規性を確立する。 2つのシミュレーション研究で手法を示す。1つはパラメトリックな非線形変換を用い、もう1つは実際の行動指標に基づくスプライン基底展開を用いる。SIR-StEMは偏りが小さく、信頼区間の被覆率が名目水準に近い結果を示し、他の一般的な方法を上回った。最後に限界と今後の研究の方向を述べる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Estimating regression models with nonlinear predictor transformations is challenging when data are missing, because nonlinearity typically renders the conditional distribution of the missing values intractable. Previous methods require specific nonlinear forms, such as polynomials or interactions, or rely on approximations that can induce bias. We propose a stochastic EM algorithm that uses sampling-importance resampling (SIR-StEM) to handle missing data under arbitrary nonlinear transformations of predictors, either substantively motivated or incidental, such as spline basis expansions. Unlike approaches that require linearity or closed-form conditionals, SIR-StEM only requires evaluating the complete-data likelihood up to proportionality, making it applicable across a broad class of nonlinear regression models. We construct an algorithm that makes use of missing data pattern information for computational efficiency, and establish asymptotic normality of the estimator. We demonstrate the method in two simulation studies, one with parametric nonlinear transformations and another with a spline basis expansion based on real behavioral measures. Results show that SIR-StEM yields low bias and near-nominal confidence interval coverage, outperforming other common approaches. We conclude with limitations and directions for future research.

著者のコメント

23 pages, 7 figures, 1 table

arXiv ID: 2609.23747 / 要約の誤りについて