観察データの因果推定で検出力計算のずれを調べる
Beyond the Illusion of Power: Calibrating Quasi-Experiments in Observational IS
この論文をやさしく読む
ひとことで言うと
準実験研究の事前の検出力計算と、実際の検出力がずれる理由を大量のシミュレーションで調べた。
何に役立つ?
観察データでDiDやIVを設計するとき、標本数だけでなく脱落や導入時期、操作変数の強さを点検する助けになる。
この研究の面白いところ
約980万データセットを使い、外生的な脱落だけで検出力が約8~11ポイント下がる条件を報告した。
どこまで分かった?
数値は要旨に示されたモンテカルロ条件、とくに数百から千程度の標本サイズに関するもの。IVの結論には第一段階F値を固定する条件が付く。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
情報システム研究では、観察パネルデータから因果効果を得るため、差の差法(DiD)や操作変数法(IV)などの準実験的方法がますます使われている。これらの設計を正当化する検出力計算は、誤差の独立同分布を仮定する。しかし、より深い問題は、クラスターに頑健な計算器にも見えない要因にある。本研究は9,837のパラメータ条件、約980万のデータセットを対象とするモンテカルロ研究を報告し、計画時と実際に達成した検出力の差を分解する。系列相関による部分は、相関係数ρが既知ならAR(1)を考慮する計算器で回復でき、短い介入前期間からρを推定する必要がある場合には一部だけ回復できる。一方、パネルからの脱落、導入時期がずれることによるバイアス、平行トレンドの事前検定は、いずれも閉じた形の公式では捉えられない。外生的な脱落だけでも、情報システム研究で使われる数百から千程度の標本サイズでは、検出力を約8~11ポイント下げる。処置と相関し、結果に依存する脱落は、検出力の低下だけでなくバイアスを生む。操作変数法では、第一段階のF値を固定すると、標本数Nを増やしても検出力は上がらず、除外制約違反によるバイアスも抑えられない。ただし操作変数自体を固定すれば、データを増やすことで第一段階の推定は精密になる。したがって識別を支えるのは標本サイズではなく操作変数の強さである。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Information systems (IS) researchers increasingly use quasi-experimental methods such as difference-in-differences (DiD) and instrumental variables (IV) to recover causal effects from observational panel data. Power calculations that justify these designs assume i.i.d. errors, but the deeper problem is what even a cluster-robust calculator cannot see. We report a Monte Carlo study over 9837 parameter conditions (approx 9.8 million datasets) and decompose the planned-versus-achieved power gap. The serial-correlation component is recoverable by an AR(1)-aware calculator when rho is known, and partially when rho must be estimated from short pre-periods, but panel attrition, staggered-adoption bias, and parallel-trends pretesting are captured by no closed-form formula; exogenous attrition alone costs approx 8 to 11 percentage points at the few-hundred-to-thousand sample sizes IS studies use. Treatment-correlated, outcome-dependent attrition instead induces bias, not just power loss. For IV, holding first-stage F fixed, larger N neither raises power nor curbs exclusion bias, though with a fixed instrument more data does sharpen the first stage, so identification rests on instrument strength, not sample size.
著者のコメント
Accepted for publication at the 60th Hawaii International Conference on System Sciences (HICSS-60)
arXiv ID: 2609.27299 / 要約の誤りについて