arXiv論文メモ
新着一覧
stat.ME / math.ST / stat.TH · 査読状況未確認

仮説間で証拠を共有する逐次的な多重検定

Pooling Sequential Evidence Across Hypotheses: Rate-Optimal Multiple Testing at a Fixed Horizon

Prasanjit Dubey and Xiaoming Huo

この論文をやさしく読む

ひとことで言うと

複数の仮説を順番に検定するとき、系列間の証拠を共有し、誤判定の確率を抑えながら早く判断するための統計手法。

何に役立つ?

限られた観測期間で複数の候補を評価する試験や実験を設計する際の理論的な選択肢になる。要旨の患者転帰に関する31%という数値は、41通りのシミュレーション設定での比較である。

この研究の面白いところ

誤った仮説が何個あるかを先に決めず、部分集合の尤度比を混合して証拠を共有する。有限期間での誤り率制御と、条件付きでの証拠の最適な成長率の両方を論じている。

どこまで分かった?

理論結果は独立系列、共通の単純な帰無・対立分布などのモデル条件に基づく。モンテカルロによる境界較正には信頼度に関する条件が付き、個々の仮説の期限内の検出力は最良の単一系列を超えない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本研究は、観測に費用がかかり、採取期間があらかじめ決まっている場合に、固定された複数の仮説を逐次検定する方法を調べる。誤った仮説の数もどれかも分からない中で、どれか一つでも誤って棄却する確率を水準α以下に抑えつつ、早い段階の判断に向けて証拠を統合することが課題となる。従来の統合法は、誤った仮説が一つの場合は平均、すべての場合は積を使うなど、特定の誤仮説数に対する証拠の成長率を実現していた。 共通の単純な帰無分布と対立分布を持つ、互いに独立した観測系列のモデルの下で、著者らは各交差仮説を、周辺尤度比の積を事前重みで混合した量によって検定する。閉検定では、この基本対称多項式で表される混合を組み合わせ、個々の誤った仮説を特定する。設計に合わせた境界の較正により、有限の観測期間で家族内誤り率を制御する。有限状態の場合には厳密な保証が得られ、モンテカルロ較正の場合には信頼度に関する条件が付く。事前分布に一致する混合は、各期間における対数証拠の期待値を一意に最大化する。空でない観測系列の部分集合すべてに正の重みを与える混合では、l個の系列が対立分布に従うとき、対数証拠の成長率はlDに達する。ここでDは対立分布下の1観測当たりの平均対数尤度比であり、1回の観測では各系列から1件ずつ得る。この成長率は、次元、仮説の真偽の構成、重みを固定し、十分長い観測期間の下でαを0に近づけるとき、交差仮説を判定するまでの遅れに関する一次の下界に達する。 所定の期限で個々の仮説について得られる検出力は、最良の単一系列の検出力を超えないが、すべての仮説が誤っている場合には閉検定が多重性による不利を取り除く。ガウス分布、バスケット型臨床試験、言語モデル、広告に関する検討で両面を例示する。主要なバスケット型試験の判定境界は1/αより40〜52%低かった。41通りのシミュレーション設定では、事前に中間解析時点を決めたボンフェローニ検定と比べ、上限を設けた平均患者転帰の中央値が31%減少した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We study sequential testing of a fixed family of hypotheses when observations are costly and a sampling horizon is specified in advance. The challenge is to pool evidence for earlier decisions when the number and identities of false hypotheses are unknown, while controlling the probability of any false rejection at level $\alpha$. Existing merges attain the pooled growth rate at a single number of false hypotheses: averaging when one is false, multiplying when all are. Under an independent-stream model with common simple null and alternative distributions, we test each intersection with a prior-weighted mixture of products of marginal likelihood ratios. Closed testing combines these elementary-symmetric-polynomial mixtures to identify individual false hypotheses. Design-specific boundary calibration gives finite-horizon family-wise error control, with exact finite-state guarantees or a confidence qualification for Monte Carlo calibration. The prior-matched mixture uniquely maximizes expected log evidence at each horizon. Mixtures assigning positive weight to every nonempty subset of streams attain log-growth rate $lD$ when $l$ streams follow the alternative. Here $D$ is the mean log likelihood ratio per alternative observation, and a round supplies one observation per stream. This rate attains the first-order intersection-delay lower bound as $\alpha\downarrow0$ at fixed dimension, configuration, weights, and a long enough horizon. Power for an individual hypothesis cannot exceed the best single-stream power at a given deadline, but closure removes the multiplicity penalty when all are false. Gaussian, basket-trial, language-model and advertising studies illustrate both. The primary basket boundaries are 40-52% below $1/\alpha$. Across 41 simulated configurations, the median reduction in capped mean patient outcomes relative to prespecified interim-look Bonferroni tests is 31%.

arXiv ID: 2609.27465 / 要約の誤りについて