成長曲線モデルで変数選択後の推論を正しく行う
Selective Inference in Growth Curve Models
この論文をやさしく読む
ひとことで言うと
同じデータで説明変数を選び、そのまま効果を検定すると推論が偏り得る問題に対処します。成長曲線モデルで、ガウス雑音を足し引きして選択用と推論用の情報を分けます。
何に役立つ?
心理学などで時間に伴う変化を調べ、どの初期特性が個人差と関連するかを選ぶ分析に役立ちます。参加者を二群に分ける方法より情報を活用しながら、選択後の推論の妥当性を保つことを狙います。
この研究の面白いところ
人を分割せず全参加者を両段階に残し、情報の方を分離します。選択された作業モデル内の共分散で重み付けした線形射影パラメータを推論対象にする点も明確です。
どこまで分かった?
共分散構造が既知なら厳密な推論、推定する場合は適切な正則条件下で漸近的な妥当性を示しています。シミュレーションと米国の若者の縦断研究への適用であり、あらゆるモデルの誤指定に無条件で対応する主張ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
成長曲線モデルは心理学研究で広く使われ、変数選択によって、時間経過に伴う異質性と関連するベースライン特性を特定できる。しかし、データに基づく変数選択の後に従来の推論を行うと、同じ結果データを選択と推論の両方に使うため、推論が妥当でなくなる可能性がある。 本研究では、ガウスノイズの加算と減算によって、すべての参加者を両段階に残しながら、選択に使う情報と推論に使う情報を分離する、成長曲線モデル用の選択後推論のデータ・フィッション枠組みを開発する。この枠組みは柔軟な変数選択手続きを扱い、選択された作業モデルにおける共分散重み付き線形射影パラメータを対象とする。共分散構造が既知なら厳密な推論を提供し、共分散構造を推定する場合にも、適切な正則性条件の下で漸近的に妥当であることを示す。 シミュレーションでは、提案法が妥当な推論を提供し、個人単位のデータ分割より効率を高める一方、素朴な選択後推論には偏りが生じることを示した。Longitudinal Study of American Youthへの適用例によって、提案枠組みを具体的に示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Growth curve models are widely used in psychological research, and variable selection can help identify baseline characteristics associated with longitudinal heterogeneity. However, conventional inference after data-driven variable selection can be invalid because the same outcome data are used for both selection and inference. We develop a data-fission framework for post-selection inference in growth-curve models that separates the information used for selection and inference while retaining all participants in both stages through the addition and subtraction of Gaussian noise. The framework accommodates flexible variable-selection procedures and targets covariance-weighted linear projection parameters in the selected working model. It provides exact inference when the covariance structure is known, and we establish asymptotic validity under suitable regularity conditions when the covariance structure is estimated. Simulations show that the proposed method provides valid inference and improves efficiency over subject-level data splitting, whereas naïve post-selection inference can be biased. An application to the Longitudinal Study of American Youth illustrates the proposed framework.
arXiv ID: 2609.19573 / 要約の誤りについて