arXiv論文メモ
新着一覧
stat.ME · 査読状況未確認

潜在変数と選択バイアスを考慮する因果関係の推定

Statistical Inference for Causal Discovery under Selection and Latent Variables via Single-Target Interventions

Xiaotian Hou, Kwangmoon Park, Hongzhe Li

この論文をやさしく読む

ひとことで言うと

観測されない要因やデータの選ばれ方による偏りがあっても、1変数ずつの介入から因果関係を推定する統計手法です。

何に役立つ?

考えられる用途は、観察・介入データから因果構造を推定する研究です。要旨ではPerturb-seqデータへの適用を例示しています。

この研究の面白いところ

各観測変数への介入の十分性と最悪の場合の必要性を理論的に示し、検定回数を変数の数の二乗に抑えています。

どこまで分かった?

誤りの制御は漸近的で、第一段階の検出力が十分であることを条件とします。要旨のデータ解析は手法の例示であり、広い生物学的結論や実験での因果効果の確証は述べていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

観察データと介入データから因果関係を見つける際、観測されない交絡要因や選択バイアスがあると、観測変数だけの有向非巡回グラフでは因果構造を十分に表せない。既存のモデルに依存しない手法は、条件付き独立性の検定を指数関数的な回数だけ要することが多く、高次元の場合に不確実性の評価も限られる。本研究は、単一の対象への介入を使い、潜在変数と選択の影響がある状況で因果関係を発見するための、特定のモデルを仮定しない統計的推論の枠組みを開発する。 文脈変数を考慮しながらシステム変数間の因果関係を捉えるため、system-induced subgraph(SIS)を導入する。最大祖先グラフ(MAG)を通じてSISの識別可能性を示し、観測された各システム変数に対する介入が一意の識別に十分であり、最悪の場合には必要でもあることを示す。その上で、第一段階の検出力が十分という条件の下、漸近的にファミリーワイズエラーを制御する二段階のグラフ推論法を構築する。観測されるシステム変数がd_X個なら、必要な統計検定は最大で(5/2)d_Xの二乗回で、各段階内で並列に実行できる。制約照会の複雑さも定数倍を除いて最適である。 この枠組みはソフト介入を扱え、パラメトリックな構造方程式の仮定を必要としない。インターフェロンβで刺激したA549肺がん細胞株のPerturb-seqデータの解析を通じて手法を例示する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Causal discovery from observational and interventional data becomes challenging in the presence of latent confounding and selection bias, where causal structure is no longer adequately represented by directed acyclic graphs over observed variables. Existing model-free methods often rely on an exponential number of conditional independence tests and provide limited uncertainty quantification in high-dimensional settings. We develop a model-free and constraint-query optimal statistical inference framework for causal discovery under latent variables and selection using single-target interventions. We introduce the system-induced subgraph (SIS) to capture the causal relations among system variables while accounting for context variables. We establish its identifiability through maximal ancestral graphs (MAGs), and show that interventions on each observed system variable are sufficient for unique identification and necessary in the worst case. Building on these results, we develop a two-stage graph inference procedure with asymptotic family-wise error control under sufficient first-stage power. For $d_X$ observed system variables, the procedure requires at most $\frac{5}{2}d_X^2$ statistical tests, parallelizable within each stage, and achieves optimal constraint-query complexity up to a constant factor. The framework accommodates soft interventions and avoids parametric structural equation assumptions. We illustrate the methods through analysis of Perturb-seq data from interferon-$\beta$-stimulated A549 lung cancer cell lines.

arXiv ID: 2609.28856 / 要約の誤りについて