AI判定と人の確認を組み合わせた低費用の仮説検定
Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation
この論文をやさしく読む
ひとことで言うと
AIの安い判定と人の確認を使い分け、誤判定の確率を制御しながら費用を抑える仮説検定法です。
何に役立つ?
AIによる品質判定を統計的な検定に使う場合、どの対象を人に確認してもらうかを決める理論的な方法になります。
この研究の面白いところ
有限標本での検定の妥当性を保ち、目標誤り率が小さくなると理論上の最小費用に一次の精度で近づきます。
どこまで分かった?
隠れた二値ラベルを持つ対象と、固定した対象プールを前提とします。費用節約の報告は数値的な評価で、具体的な現場導入の結果は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルは、出力の評価やデータへのラベル付け、システムが望む品質基準を満たすかの判定で、費用の安い評価者として使われることが増えている。しかし、AIの判断を正式な統計的推論に用いることは、それを真のラベルとして扱うだけの場合と根本的に異なる。AIの評価には偏りや雑音があり得て、厳密な仮説検定には第一種・第二種の誤りを明示的に制御する必要がある。本研究は、AIの判断と選択的な人による検証を組み合わせ、最小費用で妥当な仮説検定を行う方法を研究する。隠れた二値ラベルを持つ対象集団を考え、固定した対象の集まりを選んだ後、意思決定者は対象を選んでAIに問い合わせるか、直接人へ送るか、AIの報告を見てから人へ引き上げるか、証拠が十分になった時点で止めることができる。 指定した検定の誤り率を達成する最小費用を捉える情報理論的な下界を導き、報告結果に依存する情報フロンティアを通じて、AI情報と人による検証の価値を特徴づける。これを基に、選択的なAI評価と適応的な人への引き上げを組み合わせる、逐次的で費用を考慮した方策SCALEを開発した。SCALEは有限標本でも妥当であり、目標の誤り確率がゼロへ近づくとき、一次の項で下界に一致する。さらに、AIの出力モデルが未知の場合へ、AIと人の対になった予備データを使って拡張する。数値的な評価では、一方の情報源が明らかに優れるとSCALEは人だけ、またはAIだけの検定に近づき、安価なAI判定と選択的な人の検証の双方が有益な場合に最大の費用節約を達成する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, the decision maker can selectively query AI, send an item directly to a human, escalate an AI-scored item to a human after observing the AI report, or stop once sufficient evidence has accumulated. We derive an information-theoretic lower bound that captures the minimum cost of achieving prescribed testing errors and characterizes the value of AI information and human verification through a report-dependent information frontier. Motivated by this characterization, we develop SCALE, a sequential cost-aware policy that combines selective AI scoring with adaptive human escalation. SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. We further extend the framework to an unknown AI-output model using paired AI-human pilot data. Numerically, SCALE approaches Human-only or AI-only testing when one source clearly dominates, while achieving its largest savings when inexpensive AI judgments and selective human verification are both valuable.
arXiv ID: 2609.28859 / 要約の誤りについて