arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

初めて組む専門家の得意分野を推定しAIから判断を委ねる

RACER: Role-Aligned Competence Estimation for Human-AI Routing

Joshua Strong, Emma Sun, Alexander Capstick, Pramit Saha, Cheng Ouyang, J. Alison Noble

この論文をやさしく読む

ひとことで言うと

専門家の過去の少数の判断から、その人が今回の問題を得意とするかを推定し、AIが判断を任せる相手として適切かを決める手法です。

何に役立つ?

人間に任せられる件数が限られる状況で、AIと専門家の分担を考えるための方法になります。胸部X線ベンチマークも評価していますが、医療現場での運用効果を直接示したものとは区別が必要です。

この研究の面白いところ

クラス名そのものを手掛かりにする近道を避けつつ、同じクラス内でも症例ごとの得意不得意を推定します。ラベルを整合的に付け替えても結果が変わらない性質を理論面でも扱っています。

どこまで分かった?

PathMNISTやCIFAR-100の一部評価には模擬専門家を用いています。また、較正の良さは指標とデータセットで異なり、すべての評価で一様に優れるという報告ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

判断の委譲を学習する問題では、予測システムがいつ自律的に判断し、いつ人間の専門家に委ねるべきかを問う。専門家集団に適応する委譲は、専門家の行動に関する少量の文脈集合を用いて、この問題を未見の専門家へと拡張する。L2D-Popのようなニューラル文脈エンコーダーは問い合わせに依存できるが、絶対的なクラス座標に結び付いた振り分けの近道を学習する場合がある。Identity-Free Deferral(IFD)は、役割で索引付けしたクラスごとの能力プロファイルによってそのような近道を取り除く。しかし、その推定値は各クラス内で一定であり、個々の事例に応じた専門家の専門性を捉えられない。 私たちはRACER、すなわち振り分けのための役割整合型能力推定を提案する。これは文脈から未見の専門家の能力を推定する、役割を相対的に扱う枠組みである。RACERは、各候補クラスの役割の下で、専門家が問い合わせに正答する事後予測確率を推定し、それらをモデルの事後分布と組み合わせて、ベイズ的な判断に必要な専門家の正答確率を得る。ノンパラメトリックおよびニューラルなカーネルプーリング推定器は、候補の役割の関係、共有された集約、対称的な要約を用い、絶対的なクラス識別情報の経路を排除する。整合的なクラスのラベル付け替えに対する不変性を証明し、ベイズ基準に沿った委譲用の代理目的関数を導く。さらに、振り分けのリグレットと分類器および能力推定の誤差を関係付ける、プラグイン推定のリグレット上界を与える。 模擬専門家を用いたPathMNIST病理組織画像での文脈量を変える研究を含む、制御された合成ベンチマークでは、隠れたサブタイプへの依存があるとき、RACERは文脈の追加から恩恵を得る。また、CIFAR-100の合成実験で、別途標本化した未見の専門家の分割において、最も高い総合性能を示す。放射線科医および人間・AIの胸部X線画像ベンチマークであるVinDr-CXRとCheXpertでは、RACER系の手法は委譲予算を変えた評価で競争力のある、または最良の性能を示す。一方、確率の較正結果は指標やデータセットによって異なる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Learning to defer asks a predictive system when to act autonomously and when to defer to a human expert. Population-adaptive deferral extends this problem to unseen experts using a small context set of expert behavior. Neural context encoders such as L2D-Pop can be query-dependent, but may learn routing shortcuts tied to absolute class coordinates. Identity-Free Deferral (IFD) removes such shortcuts through role-indexed classwise competence profiles, but its estimates are constant within each class and cannot capture instance-level expert specialization. We propose RACER---Role-Aligned Competence Estimation for Routing---a role-relative framework for estimating an unseen expert's competence from context. RACER estimates the posterior-predictive probability that the expert is correct on a query under each candidate class role, then combines these estimates with the model posterior to obtain the Bayes-relevant expert-correctness probability. Nonparametric and neural kernel-pooling estimators use candidate-role relations, shared aggregation, and symmetric summaries, excluding absolute class-identity channels. We prove coherent class-relabelling invariance, derive a Bayes-aligned deferral surrogate, and give a plug-in regret bound relating routing regret to classifier and competence-estimation error. On controlled synthetic benchmarks, including a PathMNIST histopathology context-scaling study with simulated experts, RACER benefits from additional context under hidden subtype dependence and gives the strongest aggregate performance on a separately sampled unseen-expert split in the CIFAR-100 synthetic experiments. On the radiologist and human--AI chest-radiography benchmarks (VinDr-CXR and CheXpert), the RACER family is competitive or best in budget-swept deferral, with calibration results varying across metrics and datasets.

arXiv ID: 2609.21953 / 要約の誤りについて