arXiv論文メモ
新着一覧
stat.ME / cs.IT / math.IT / math.ST / stat.TH · 査読状況未確認

分散した多重検定で報告の細かさが検出力に与える影響

Report resolution in federated multiple testing under family-wise error control

Prasanjit Dubey and Xiaoming Huo

この論文をやさしく読む

ひとことで言うと

複数機関が生データを共有せず、p値を粗い報告へ圧縮して共同検定すると、どれだけ検出力が失われるかを調べています。

何に役立つ?

通信量や報告の細かさと統計的な検出力の釣り合いを設計する理論になります。仮説群全体での誤り率を管理したまま、各機関の報告を結合する規則も扱います。

この研究の面白いところ

情報損失の原因を分散そのものではなく圧縮に分け、条件のもとで最適な損失が報告の分解能の二乗の逆数で減ることを示します。例では3ビット報告で集中処理の検出力の99.2%を保持します。

どこまで分かった?

損失率の定理には正則性と検出力の目的に関する条件があります。99.2%は有意水準0.05、2つのBeta(1,2)仮説という例であり、どの分布でも同じ性能になる保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

複数の機関が、ファミリーワイズ誤り率を制御しながら一つの仮説族を検定するが、データを統合できないため、各拠点は各仮説について自拠点のp値の報告だけを公開する。その報告が取り得る値の数を、報告の分解能と呼ぶ。報告方式と結合規則を最適に選んだ場合について、最も検出力の高い集中型の手続き(オラクル)と比べ、こうした報告によってどれだけ検出力が失われるかを定量化する。報告からp値を復元できれば検出力は失われないため、損失の原因は分散化ではなく圧縮にある。 明示した正則性条件と検出力の目的関数に関する条件の下では、どの有限分解能でも検出力は低下し、最適な損失は分解能の逆二乗で減衰する。等幅区間はこの次数を達成し、同じ数のセルへのどの分割も、この次数を改善できない。二拠点のベータ分布の例では、最適化した1ビット報告は一つの仮説に対してはほぼ損失がないが、仮説が二つになると数倍の検出力損失が生じる。二つのBeta(1,2)仮説について、有意水準0.05では、最適に結合した3ビットの等幅報告がオラクルの検出力の99.2%を保持する。 与えられた報告に対して検出力が最適となる結合規則と、報告の正確に既知の帰無分布だけを使い、仮説間に任意の依存関係があっても有効な規則を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Several institutions test one family of hypotheses under family-wise error control but cannot pool their data, so each site releases, for each hypothesis, only a report of its own p-value. The report's resolution is the number of values it can take. We quantify the power lost to such reports relative to the most powerful centralized procedure (the oracle), with reports and combining rule chosen optimally. A report from which the p-value can be recovered loses no power, so the loss is due to compression, not decentralization. Under the stated regularity and power-objective conditions, every finite resolution loses power, and the optimal loss decays as the inverse square of the resolution: equal-width intervals attain this order, and no partition into as many cells improves it. In two-site Beta examples, optimized one-bit reports are nearly lossless for a single hypothesis but lose several times as much power for two. At level 0.05 with two Beta(1, 2) hypotheses, optimally combined three-bit equal-width reports retain 99.2% of oracle power. We give the power-optimal combining rule for given reports, and a rule using only their exactly known null distribution, valid under arbitrary dependence across hypotheses.

arXiv ID: 2609.19708 / 要約の誤りについて