arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

安全な予測を何件まで提供できるかを計画する

Available Guardrails: Certifying Selective Prediction across ML Systems

Parivesh Priye, Yufeng Wang, Haibin Ling, Michael Chaykowsky

この論文をやさしく読む

ひとことで言うと

信頼できる予測だけを返す仕組みで、安全性の基準を満たす証拠が足りるかと、どれだけの利用者に結果を返せるかを一緒に計画します。

何に役立つ?

ツールや利用者群ごとに目標適合率を保証したい運用で、細かい区分と回答できる件数の折り合いを考える手掛かりになります。

この研究の面白いところ

安全性の保証が正しいかだけでなく、データ不足で保証そのものを出せない問題を扱います。区分を作るデータと選ぶデータを分ける改善策も評価しています。

どこまで分かった?

区分選択は群の順序を固定した条件です。0.157の改善は真の情報を知る計画器の値で、実際の有限データで得た改善0.060と区別する必要があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

選択的予測器は、安全性のゲートとして働き、予測が十分に信頼できると見える場合にだけ出力する。実運用では、ツール、ポリシーラベル、患者の部分集団など、関心のある各報告単位について、目標の適合率でこの信頼性を認証することがますます求められる。主な難しさは、与えられた認証が有効かどうかより、有限の較正データからそもそも認証を得られるかどうかであることが多い。ゲートをより安全に、またはより細かくすると、一部の単位では認証に足る証拠が少なすぎる場合がある。 本研究では、古典的な厳密二項分布の反転により、この提供可能性を計算可能にする。また、群の順序を固定した下での報告区分の選択を動的計画問題として定式化し、安全性、細かさ、処理できるトラフィックの間のトレードオフを明らかにする。得られたフロンティアは、母集団では大きい改善機会が、有限標本での推定によってほぼ消えてしまうことを示す。真の情報を知る計画器は、各群の支持数を均衡させる方式より平均カバレッジを0.157改善する一方、単純な推定器が回復するのは0.005だけであり、有限データからの回復が中心的課題となる。 一つの計画用分割で候補区分を構築し、別の分割で候補を選ぶと、この差の一部を回復でき、支持数を均衡させる方式に対して平均カバレッジが0.060改善する。この改善方向は、3つの意図振り分けデータセットと2つの構造にわたる60のモデル効果のうち59で再現された。有効性を保つ相補的な手段として、報告単位間でファミリーワイズ誤りの予算を再配分すると、母集団の量と雑音を含む推定値の両方で、さらにカバレッジを回復できる。同じフロンティアは、予測器ごとの上限を伴いながら、LLMのツール呼び出し、コンテンツモデレーション、病変分類、推薦にも現れる。したがって、認証された提供可能性は、いつ、どの細かさで、どれだけのトラフィックについて安全ゲートを認証できるかを決める、計画可能な運用資源である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup. The main difficulty is often not whether a granted certificate is valid, but whether finite calibration data can produce one at all. As the gate becomes safer or more fine-grained, some units may receive too little evidence to certify. We make this notion of availability computable through classical exact-binomial inversion and formulate reporting-partition selection, under a fixed group order, as a dynamic program that exposes the trade-off among safety, granularity, and served traffic. The resulting frontier reveals a large population opportunity that finite-sample estimation nearly erases: a truth-informed planner gains $0.157$ mean coverage over support balancing, whereas a naive estimator recovers only $0.005$, making recovery from finite data the central challenge. Constructing candidate partitions on one planning split and selecting among them on another recovers part of this gap, improving mean coverage over support balancing by $0.060$, with the direction reproduced in $59$ of $60$ model effects across three intent-routing datasets and two architectures. A complementary validity-preserving lever, reallocating the familywise error budget across reporting units, recovers additional coverage both with population quantities and noisy estimates. The same frontier recurs, with predictor-specific ceilings, across LLM tool-calling, content moderation, lesion classification, and recommendation. Certified availability is therefore a plannable deployment resource that determines when a safety gate can be certified, at what granularity, and over how much traffic.

arXiv ID: 2609.22048 / 要約の誤りについて