arXiv論文メモ
新着一覧
stat.ME · 査読状況未確認

多段階標本抽出で最良または最悪の集団を選ぶ方法

Extreme Population Selection under Multistage Sampling design With Applications

Shivam, Bhargab Chattopadhyay, Nil Kamal Hazra

この論文をやさしく読む

ひとことで言うと

複数の集団から指標が最良または最悪の集団を、段階的に標本を集めながら選ぶ二つの方法を提案した。

何に役立つ?

考えられる用途は、不平等の指標や腫瘍変異量のような推定値を使い、調査対象の集団を信頼度を指定して選ぶこと。

この研究の面白いところ

分布の形を固定する仮定を置かずに、オンライン方式と多腕バンディット方式の両方について選択の正しさを示した。

どこまで分かった?

理論上は集団間に十分な差と正則性条件が必要。応用例は経済分野のシミュレーションと指定の臨床シーケンス集団に基づく。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

K個(K≥2)の集団の中から最良または最悪の集団を選ぶ問題を、極端な集団が次点の集団から十分に離れていると仮定して調べる。選択には適切な指標を使い、その指標は応用分野によって異なりうる。指標の真の値は不明なので、多段階標本抽出の設計の下で、一般化モーメント法により推定量を求める。この推定量を用い、オンラインアルゴリズムと多腕バンディットに基づくアルゴリズムという二つの逐次的な手順を提案する。適切な正則性条件の下で、基礎となる分布にパラメトリックな仮定を置かず、両手順は指定した信頼水準で極端な集団を正しく特定する。 提案手法を計量経済学と遺伝学の応用で示す。計量経済学では、不平等の指標であるジニ係数を用いて極端な集団を選び、さまざまな分布の設定で幅広いモンテカルロシミュレーションを行い、手順の性能を評価した。遺伝学では、腫瘍変異量(TMB)スコアから導いた指標で最悪の集団を特定し、Memorial Sloan Kettering-IMPACT 50000臨床シーケンス集団を用いて実用上の適用可能性を示した。さらに、異常な集団が存在する場合には、その集団を特定するためにも提案した枠組みを用いる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We study the problem of selecting the extreme (best or worst) population from among $K(\geq 2)$ populations, under the assumption that the extreme population is sufficiently separated from the nearest population. The selection is based on an appropriate measure, which may vary across different application domains. Since the actual value of the measure is unknown, we obtain its estimator using the generalized method of moments under a multistage sampling design. Using this estimator, we propose two sequential procedures, namely online algorithm and multi armed bandit based algorithm. Under suitable regularity conditions and without imposing parametric assumptions on the underlying distributions, both algorithms correctly identify the extreme population with a desired level of confidence. We illustrate the proposed algorithms through applications in econometrics and genetics. In the econometric application, the extreme population is selected using the Gini index as a measure of inequality and the performance of the proposed procedures is assessed through extensive Monte Carlo simulation studies conducted under various distributional settings. In the genetics application, the worst population is identified using a measure derived from the tumor mutation burden (TMB) score and the practical applicability of the proposed algorithms is demonstrated using the Memorial Sloan Kettering-IMPACT 50000 clinical sequencing cohort. Further, we use the proposed framework to identify an anomalous population, provided such a population exists.

arXiv ID: 2609.28127 / 要約の誤りについて