arXiv論文メモ
新着一覧
stat.ME / math.ST / stat.ML / stat.TH · 査読状況未確認

必要な変数が丸ごと欠けた集団で回帰を推定する

Doubly robust target inference for generalized linear regression with completely missing covariates

Huali Zhao (1), Ke Deng (2) ((1) School of Mathematics and Statistics, Huazhong University of Science and Technology, (2) Department of Statistics and Data Science, Tsinghua University)

この論文をやさしく読む

ひとことで言うと

ある集団で必要な項目を一人分も測っていなくても、その項目を測った別の集団のデータを使い、一定の仮定の下で回帰を推定する方法です。

何に役立つ?

既存のコホートやバイオバンクにない変数を含む分析を考える際に、何を仮定すれば推論できるかを整理する理論的基盤になります。

この研究の面白いところ

二つの補助モデルのどちらかが正しければ一致性を保つ二重頑健性を、非線形な一般化線形回帰で扱います。単純な平均の補完だけでは足りない点に対応しています。

どこまで分かった?

観測変数を与えたときの欠測変数の条件付き分布が集団間で共通という仮定が必要です。二重頑健性がこの仮定を不要にするわけではなく、効率性の下界達成には両モデルの正しい指定も必要です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模で多目的なコホート研究やバイオバンクでは、特定の後続分析に必要な共変量が収集されていないことが多い。本研究では、重要な共変量が対象集団のデータには完全に存在せず、関連する情報源集団では観測されている場合に、一般化線形回帰について対象集団での推論を行う問題を扱う。標準的な共変量欠測への対処法は、対象集団で少なくとも一部の共変量が観測されていることを必要とするため、そのままでは適用できない。 本研究は、部分集団シフトの仮定の下で、二重頑健な転移学習の枠組みを開発する。この仮定は、観測された結果変数と共変量の分布が集団間で異なることを許す一方、観測変数を条件とした欠測共変量の条件付き分布が共通であることを要求する。線形回帰と異なり、非線形な対象集団の推定方程式には、固定された少数の低次条件付きモーメントではなく、パラメータで添字付けされた欠測共変量の条件付き汎関数が必要になる。 提案する推定量は、重要度で重み付けした情報源集団の推定方程式と、これらの条件付き汎関数の補完を組み合わせる。識別のための仮定の下では、密度比モデルまたは条件付き共変量モデルのどちらかが正しく指定されていれば、推定量は一致性を保つ。正則条件の下では√n一致性と漸近正規性を持ち、両方の補助モデルが正しく指定されている場合には、セミパラメトリックな効率性の下界を達成する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large-scale multipurpose cohort studies and biobanks often omit covariates needed for specific downstream analyses. We study target-population inference for generalized linear regression when key covariates are completely absent from the target data but observed in a related source population. Standard missing covariate methods are not directly applicable because they require at least partial observation of the covariates in the target population. We develop a doubly robust transfer learning framework under a sub-population shift assumption, which allows the distribution of the observed outcome and covariates to differ between populations while requiring the conditional distribution of the missing covariates given the observed variables to be shared. Unlike linear regression, nonlinear target estimating equations require parameter-indexed conditional functionals of the missing covariates rather than a fixed collection of low-order conditional moments. Our estimator combines importance-weighted source estimating equations with imputation of these conditional functionals. Under the identifying assumption, the estimator remains consistent when either the density-ratio model or the conditional-covariate model is correctly specified. Under regularity conditions, it is root-$n$ consistent and asymptotically normal, and attains the semiparametric efficiency bound when both nuisance models are correctly specified.

著者のコメント

22 pages, 5 tables

arXiv ID: 2609.24086 / 要約の誤りについて