arXiv論文メモ
新着一覧
cs.CL / cs.LG · 査読状況未確認

少数の正解ラベルで言語モデルの順序尺度を校正

CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels

Xiangwei Wang, Peng Wang, Saman Halgamuge

この論文をやさしく読む

ひとことで言うと

言語モデルが出す段階評価の確率を、少数の正解例で補正する方法です。

何に役立つ?

レビューや会話の評価で、ラベルを大量に集めにくい場合の確率校正に役立つと考えられます。

この研究の面白いところ

5つの解釈可能なパラメータを使い、80設定中76設定で比較法の中で最小の対数損失を得ました。

どこまで分かった?

優位性は主に5~100件の少数ラベルで示されました。数百~数千件では制約の少ない校正法が上回る場合があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルは文章から順序のある尺度の確率分布を出せるが、その分布は、値が端に寄る、圧縮される、誇張される、ある方向へ一貫して偏るなどのノイズを含む測定値である。本研究は、モデルの出力を真のラベルのノイズを含む読み取りとみなし、解釈しやすい5つのパラメータを持つチャネルで補正するCORDIALを提案する。このチャネルは小さいため、少数のラベルから得た事後分布を平均できる。また、この校正が一次確率順序を保つことを証明する。 AmazonのレビューとCMU-MOSEIの書き起こしに4つの言語モデルを用いた評価では、5~100件のラベルを使う80設定のうち76設定で、9種類の校正法中、対数損失が最も低かった。主に使った7Bモデルで20件のラベルを用いると、28~54件のラベルを使う最も強い比較手法に並んだ。同じ事後分布を使えば、別の課題から事前分布を学ぶことや複数の言語モデルの出力を融合することもできる。ディリクレ校正など制約の少ない手法が上回るのは、校正用データが数百から数千件に増えた場合に限られた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with a channel of five interpretable parameters. The channel is small enough for its posterior to be averaged from a handful of labels, and we prove that the resulting calibration preserves first-order stochastic order. On Amazon reviews and CMU-MOSEI transcripts with four LLMs, CORDIAL has the lowest log loss among nine calibrators in 76 of 80 settings with 5 to 100 labels; with 20 labels and the main 7B reader, it matches the strongest baseline using 28-54 labels. The same posterior lets us learn priors from other tasks and fuse several LLMs. Unrestricted calibrators such as Dirichlet calibration overtake it only as the calibration set grows into the hundreds or thousands.

arXiv ID: 2609.29807 / 要約の誤りについて