arXiv論文メモ
新着一覧
cs.LG / cs.CL · 査読状況未確認

追加学習後の言語モデルの過信をラベルなしで補正する

SupportCal: Label-Free Calibration of Post-Trained LLMs via Reference Support and Corroboration

Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, José Miguel Hernández-Lobato, Junbin Gao

この論文をやさしく読む

ひとことで言うと

追加学習した言語モデルが答えを信じすぎる問題を、元のモデルと他の参照モデルの判断を使って補正します。不一致の例も重みを変えて利用します。

何に役立つ?

専用の正解ラベルを集めにくい場合に、確信度を調整する方法として役立ちます。評価された較正誤差の改善と、実際の利用場面での安全性は別の問題です。

この研究の面白いところ

不一致例を増やせば必ず改善するわけではないという診断から、例ごとの連続的な重み付けを設計しています。

どこまで分かった?

改善は評価したほぼすべての構成であり、例外がないという主張ではありません。最適温度の具体的な存在条件や改善量は要旨に示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

追加学習はタスク性能を改善することが多いが、確信度の較正を悪化させ、追加学習後の言語モデル(PoLM)が対応する事前学習言語モデル(PLM)よりも過信する状態になる場合がある。タスク固有のラベル付き較正データは、入手に費用がかかるか、利用できないことがあるため、元のPLMはラベル不要の事後較正における自然な参照となる。従来の、一致を条件とするPLM参照型較正は、PoLMと参照PLMの回答が一致する例だけを用いて単一の温度パラメータを適合させる。不一致例は、直接的な整合によって適合温度が過度に高くなり、確信度が低すぎる状態を招き得るため除外していた。 本研究は、この二者択一的な扱いを再検討する。不一致例を制御しながら再導入する診断により、集計した効果が単調でないことが分かった。不一致例を適度な割合で取り込むと較正が改善する場合がある一方、重み1での採用範囲が不一致例全体に近づくにつれて利点は小さくなる。提案するSupportCalは、一致例の重みを1に保ち、不一致例には、当該モデルの元となるPLMによる相対的な支持と、規模の適合する候補群から選んだ事前学習済み参照モデルによる裏付けに基づいて連続的な重みを割り当てる、ラベル不要の事後較正手法である。さらに、得られる重み付き目的関数が有限の最適温度を持つ条件を特徴付ける。MedMCQAとMathQAでは、評価したほぼすべての対象モデル構成で、一致例だけを用いるベースラインより低い期待較正誤差(ECE)を達成した。補足的なTweetEval Sentimentの結果でも、ラベル集合が固定された分類タスクで同じ傾向が見られた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident than their corresponding pretrained language models (PLMs). Because task-specific labeled calibration data can be costly or unavailable, the corresponding pretrained PLM provides a natural label-free reference for post-hoc calibration. Prior agreement-gated PLM-referenced calibration fits a scalar temperature using only examples on which the PoLM and its PLM reference agree, excluding disagreement examples because direct alignment can drive the fitted temperature excessively high and induce under-confidence. We revisit this binary treatment. A controlled reintroduction diagnostic reveals a non-monotonic aggregate effect: admitting a moderate fraction of disagreement examples can improve calibration, whereas the benefit diminishes as unit-weight inclusion approaches the full disagreement set. We introduce SupportCal, a label-free post-hoc method that retains agreement examples at unit weight and assigns disagreement examples continuous weights based on the own-base PLM's relative support and corroboration from pretrained references selected from a size-compatible candidate pool. We further characterize when the resulting weighted objective admits a finite optimal temperature. Across MedMCQA and MathQA, SupportCal yields lower ECE than the agreement-only baseline for nearly all evaluated target-model configurations; supplementary TweetEval Sentiment results show the same pattern on a fixed-label classification task.

著者のコメント

14 pages, 5 figures, 6 tables

arXiv ID: 2609.24303 / 要約の誤りについて