複数ラベル分類の確信度のずれを測る方法
How to Estimate Whether You Have Found Several Needles in a Haystack: Measuring Calibration in Multi-Label Text Classification
この論文をやさしく読む
ひとことで言うと
一つの文書に複数のラベルを付ける分類で、確信度が正しさに合っているかを測る方法。
何に役立つ?
医療コードなど負例が非常に多い分類モデルの確信度評価に役立つ。
この研究の面白いところ
正例と負例に等しい重みを置く区間分けで、従来指標の過小評価やラベル頻度への依存を避けようとする。
どこまで分かった?
提案は校正誤差の測定法であり、大規模言語モデルの確信度を校正する問題はなお未解決としている。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自動予測を信頼できるかを決める重要な要素は確信度であり、正しい確率に見合うよう校正されているべきである。多くの確信度校正指標は二値分類や多クラス分類向けで、複数ラベル分類の校正は十分に研究されていない。診療記録への医療コード付与やニュースの話題判定のような複数ラベル分類では、該当しないラベルという負例が大量にあることが普通である。既存のラベルごとの期待校正誤差を求める区間分けは、誤差を過小評価するか、単にラベルの頻度を反映するか、事例数が非常に少ない区間を多数作ることを示す。信頼できるラベル別の校正誤差を得るため、正例と負例のラベル付与に等しい重みを与える新たな区間分けを提案する。実証研究では、従来方式と異なり、階層型および極端にラベル数が多い複数ラベル分類で意味のある校正誤差推定が得られた。また、大規模言語モデルの複数ラベル予測の確信度を校正することは未解決の課題であると示す。詳細な分析と確かな評価指標を提供し、今後の複数ラベル分類の校正研究の基盤とする。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence calibration metrics target binary or multi-class tasks, while multi-label calibration remains largely underexplored. Multi-label classification tasks, such as assigning medical codes to clinical notes or determining news topics, are usually dominated by a large number of negatives, i.e., labels that do not apply. We show that existing binning schemes to compute label-wise expected calibration error either underestimate the error, simply reflect label frequency, or suffer from many bins with very few instances. To achieve trustworthy label-wise calibration errors, we propose a new binning scheme that gives equal weight to positive and negative label assignments. Our empirical study demonstrates that in contrast to existing binning schemes, our new scheme results in meaningful estimates of calibration error in hierarchical and in extreme multi-label classification. We also show that calibrating confidence scores of large language models for multi-label predictions is an open challenge. Our detailed analysis lays the foundation for further research by providing a solid evaluation metric for measuring calibration in multi-label classification.
arXiv ID: 2609.26468 / 要約の誤りについて