無関係な変数が多くても効果の異なる集団を探す
Shrinkage Bayesian Causal Forest with Instrumental Variable
この論文をやさしく読む
ひとことで言うと
説明変数が多く、その大半が効果の違いに無関係でも、介入への反応が異なる分かりやすい集団を探します。
何に役立つ?
操作変数を使う因果分析で、遵守者の効果が誰によって変わるかを調べる用途があります。
この研究の面白いところ
ベイズ木の変数選択を疎にし、その情報を解釈しやすい別の木の分割コストにも引き継ぎます。
どこまで分かった?
真の分割の復元や被覆率の改善はモンテカルロ実験の結果です。実データ2種への適用は述べられていますが、要旨に効果量や操作変数の妥当性の詳細はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
不完全な遵守がある操作変数分析では、遵守者への効果が平均から異なる、解釈可能な部分集団を見つけることが中心的な目標である。しかし、既存の木に基づく手法は、多くの共変量が効果に無関係な場合に性能が低下する。疎な高次元設定で、異質な遵守者平均因果効果(CACE)を持つ部分集団を発見・推定する手法として、操作変数を用いた縮小ベイズ因果フォレストSBCF-IVを提案する。 SBCF-IVは、条件付きのintention-to-treat効果と遵守者割合を推定するベイズ加法回帰木の分岐確率に、疎性を促すディリクレ事前分布を置く。遵守者効果を修飾する少数の共変量へ事後確率を集中させ、効果推定を正則化する。また、事後の分岐頻度を、後段のCARTに変数ごとのコストとして入力し、関連する効果修飾変数へ分割を誘導する。これにより共変量空間を解釈可能な形で分割する。モンテカルロ実験では、無関係な共変量の割合が増えるにつれ、疎性を用いない従来のBCF-IVより、木全体と各個体の両水準で真の分割を確実に復元することを示す。さらに、BCF-IVの区間推定が悪化する場合にも、名目上の被覆率を維持する。オレゴン健康保険実験と401(k)加入資格のデータに本手法を適用する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Discovering interpretable subgroups whose complier effects deviate from the average is a central goal of instrumental variable analysis under imperfect compliance, yet existing tree-based methods degrade when most covariates are irrelevant to the effect. We propose Shrinkage Bayesian Causal Forest with Instrumental Variable (SBCF-IV) for discovering and estimating subgroups with heterogeneous Complier Average Causal Effects (CACE) in sparse high-dimensional settings. SBCF-IV places a sparsity-inducing Dirichlet prior on the splitting probabilities of the Bayesian Additive Regression Trees that estimate the conditional intention-to-treat and the complier share, concentrating posterior mass on the few covariates that moderate the complier effect and thereby regularizing effect estimation. The posterior split frequencies additionally enter a downstream CART as variable-level costs that steer the partition toward relevant moderators, providing an interpretable division of the covariate space. Monte Carlo experiments show that, as the share of irrelevant covariates grows, SBCF-IV recovers the true partition more reliably than its non-sparse predecessor BCF-IV at the tree and unit level, and retains nominal coverage where BCF-IV's intervals deteriorate. We apply the method to the Oregon Health Insurance Experiment and the 401(k) eligibility data.
著者のコメント
31 pages
arXiv ID: 2609.18903 / 要約の誤りについて