検索中心の抽出型質問応答でモデルの確信度が役立たない条件
Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute
この論文をやさしく読む
ひとことで言うと
検索だけでほぼ上限に達する文書QAで、モデル自身の確信度を使う追加学習や回答選択が役立つかを検証しています。
何に役立つ?
確信度による制御を加える前に、改善の余地があるかと、その信号が実際に有効かを点検する参考になります。
この研究の面白いところ
事前に決めた基準で4モデル系列の蒸留トリガーと単一モデルの選択・棄権方策を評価し、どちらも基準を満たさなかった結果を報告します。
どこまで分かった?
92〜99.8%は検索が最良の組合せの精度にどれだけ達するかの比率で、絶対正答率ではありません。答えを含む文章から質問を生成した抽出QAの設定で、全QAへの否定的結論ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
答えを含む文章から質問を生成した抽出型の文書質問応答では、絶対的な正解率がどの水準であっても、検索だけで、処理方式をどのように組み合わせても到達できる性能の92〜99.8%を回収できる。このような状況では、確信度を用いる仕組みが得られる改善の余地は小さい。専門分野のコーパスで公開言語モデルを微調整すると、そのモデル自身の確信度を、どの質問に追加適応が必要か、どの回答を信頼するかを判断する制御信号として利用したくなる。 本研究では、実行前に固定した基準の下で、この2つの用途を評価する。対象は70億〜90億パラメータの4モデル系列で、適応による検索なしのクローズドブックF1の変化は最大でも+0.03だった。結果はどちらも不成功である。蒸留を開始する判定は、事前に定めた3ステップの転移予算の下で4系列すべてで失敗し、処理方式の振り分けと回答棄却の方策も、単一モデルによる予備評価で失敗した。試したすべての正誤判定基準で、検索だけで最良の組合せの正解性能の92〜99.8%を回収しており、振り分けに意味のある改善余地が残らない。 系列尤度による信号は、適応の前後とも、この検索方式に対して不十分だった。登録済みの判定基準におけるROC曲線下面積は0.65〜0.81であり、スカラーによる再較正でも変わらず、トークン単位の温度再調整でも一貫した改善はなかった。さらに細かな診断結果は、正誤判定基準と回答の長さに依存する。検証できた適応済みの3つの組合せでは、選択器のアブレーションから、どの乱数シードでも確信度項による後段の便益を統計的に検出できなかった。Gemmaでは、この項を除くことで、選択器は登録済みの両基準を満たさない状態から、両方を満たす状態へ変わった。本研究で利用可能な成果は、依存条件を明示した、事前に規定された一連の否定的結果である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy -- confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields a model whose own confidence is a tempting control signal: it could decide which queries warrant further adaptation, and which answers to trust. We evaluate both uses under criteria fixed before the runs were executed, across four 7-9B model families whose adaptation moved closed-book F1 by at most +0.03, and both fail: a distillation trigger on all four families, under its pre-specified three-step transfer budget, and a routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy under every correctness criterion we test, leaving routers no meaningful gain. The sequence-likelihood signal is insufficient relative to that mode -- area under the receiver operating characteristic curve 0.65-0.81 under the registered criterion -- before adaptation as well as after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. And the finer diagnostics depend on the correctness criterion and on answer length; on the three adapted combinations where we could test it, selector ablations show no statistically detectable downstream benefit from the confidence term on any seed; on Gemma, removing it changes the selector from failing to passing both registered criteria. The usable product is a set of pre-specified negatives with their dependencies made explicit.
著者のコメント
26 pages main text + 26 pages supplementary (Online Resource 3). Submitted to Applied Intelligence. Code and data: doi:10.5281/zenodo.22710121, doi:10.5281/zenodo.22721044
arXiv ID: 2609.19942 / 要約の誤りについて