arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

言語モデルの自己申告する確信度はどこまで役立つか

An Analysis of Training-Free Self-Reported Confidence in Language Models

Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li

この論文をやさしく読む

ひとことで言うと

言語モデルが自分の答えに付ける確信度は、正誤の見分けにどこまで役立つかを調べます。

何に役立つ?

確信度表示や複数回回答の一致を使った品質判定を評価する材料です。二つのモデル系列で同じ100問を比較します。

この研究の面白いところ

自己申告の確信度は高い識別力を示した一方、三つの追加回答の一致は弱く、誤答への全員一致もありました。同じ答えでも質問の仕方で確信度が変わります。

どこまで分かった?

TriviaQAの100問と、探索的な伝記の100主張の検証です。全用途で同じ精度を保証するものではなく、ベンチマークの誤りや相関した誤答にも注意を向けています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルは生成内容とともに数値的な確信度を示せるが、それが校正された言い回し以上の意味をもつかは明らかでない。本研究では、2つのモデル系列に同一のTriviaQAの100問を与え、追加学習を必要としない3つの信号を分析した。回答とともに言語化される確信度、回答後に求めるP(True)、および同じ質問に対する追加3回の生成との一致度である。 直接言語化された確信度は、驚くほど強力な基準手法となった。ベンチマークの誤りを点検した後では、正誤予測のAUROCは0.956と0.937に達した。3サンプルの一致度は大幅に劣り、0.765と0.790だった。また、言語化された確信度との固定比率による補間にも、統計的に信頼できる改善はなかった。一方のモデルの9件の誤答のうち4件、もう一方の8件の誤答のうち2件では、すべての追加サンプルが誤答を支持した。これは、自己整合性が共通の誤解を増幅しうることを示す。 同じ固定済み回答について、同等のプロンプトで確信度を再度尋ねると、スコアは平均0.043〜0.084変化し、しきい値0.8での判定の4〜9%が反転した。さらに、確信度付きの人物略歴に関する100件の主張を探索的に点検したところ、裏付けのある主張と反証される主張の確信度の差は小さいものにとどまった。これらの結果は、自己申告値が有用であっても、尋ね方、相関した誤り、ベンチマークのノイズに影響されることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\mathrm{True})$, and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is a surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction. Three-sample agreement is substantially weaker (0.765 and 0.790), and a fixed interpolation with verbalized confidence has no statistically reliable benefit. Four of nine errors from one model and two of eight from the other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions. Re-eliciting confidence for the same fixed answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4\% to 9\% of decisions at a 0.8 threshold. An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims. These results argue that useful self-reports remain sensitive to elicitation, correlated errors, and benchmark noise.

著者のコメント

workshop

arXiv ID: 2609.20541 / 要約の誤りについて