複数の選抜条件で偏る計数データの回帰を補正する
Poisson Regression under Multivariate Sample Selection
この論文をやさしく読む
ひとことで言うと
複数の条件を通過した対象しか観測できない計数データで、観測対象の偏りを補正する回帰モデルです。
何に役立つ?
観測されるかどうかと結果が関連するデータを分析する際に、選ばれた標本だけの回帰が持つ偏りを考慮するために役立ちます。
この研究の面白いところ
選択条件を一つに限定せず複数扱い、第1段階の選択モデルの推定誤差まで共分散の計算に反映しています。
どこまで分かった?
導出は誤差の同時正規性などを仮定します。除外制約なしの識別にも台と階数の条件が必要です。性能比較はモンテカルロシミュレーションで、実データ適用は要旨に記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本論文では、相関し得る複数の選択条件が満たされた場合にのみ結果が観測される、多変量の標本選択を伴うポアソン回帰モデルを開発する。著者らの知る限り、任意の数の選択方程式を許す初めてのポアソン標本選択モデルである。結果と選択の誤差が同時正規分布に従う下で、観測される結果の条件付き平均を導出し、多変量の選択補正項を得る。 モデルのパラメータの識別を証明し、適切な台と階数の条件の下では、除外制約なしで結果側のパラメータを識別できることを示す。また、非線形最小二乗法とポアソン疑似最尤法(PPML)に基づく二段階推定手順を提案する。著者らの知る限り、標本選択を伴うポアソン回帰モデルへのPPMLの適用は本論文が初めてである。両推定量の一致性を確立し、第1段階の選択モデルの推定誤差を考慮する、頑健な二段階サンドイッチ共分散行列を提案する。さらに、階乗モーメントを用いて、潜在的な結果誤差の分散と、結果誤差・選択誤差の間の相関を復元する。 モンテカルロシミュレーションは、結果と選択の誤差が相関する場合、標本選択を無視すると持続的なバイアスが生じることを示す。一方、提案PPML推定量はこのバイアスを大幅に減らし、特に選択の依存関係が中程度から強い場合、非線形最小二乗法より安定している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
This paper develops a Poisson regression model with multivariate sample selection, in which the outcome is observed only when several potentially correlated selection conditions are satisfied. To the best of our knowledge, this is the first Poisson sample selection model that allows for an arbitrary number of selection equations. We derive the conditional mean of the observed outcome under joint normality of the outcome and selection errors and obtain a multivariate selection-correction term. We prove identification of the model parameters and show that, under suitable support and rank conditions, the outcome parameters can be identified without an exclusion restriction. We also propose two-step estimation procedures based on nonlinear least squares and Poisson pseudo-maximum likelihood. To the best of our knowledge, this paper is the first to apply PPML to a Poisson regression model with sample selection. The consistency of both estimators is established, and a robust two-step sandwich covariance matrix is proposed to account for the estimation error from the first-step selection model. In addition, factorial moments are used to recover the variance of the latent outcome error and the correlations between the outcome and selection errors. Monte Carlo simulations show that ignoring sample selection leads to persistent bias when the outcome and selection errors are correlated, while the proposed PPML estimator substantially reduces this bias and is more stable than nonlinear least squares, especially under moderate and strong selection dependence.
著者のコメント
35 pages
arXiv ID: 2609.21056 / 要約の誤りについて