人間とAIの会話を評価する言語モデル判定者の較正
Calibrating LLM Judges for Human and AI Conversations
この論文をやさしく読む
ひとことで言うと
会話を評価する言語モデルの点数を比較可能にするため、基準例を使って較正した研究。
何に役立つ?
異なる言語モデル判定者の採点を同じ尺度で比較し、会話評価のずれを調べるのに役立つ。
この研究の面白いところ
人間とAIの会話200件を含む新しいデータで、別の会話データから作った較正が移るかを検証した。
どこまで分かった?
単独採点と人間評価の相関は中程度で、対比較には長文と提示位置の影響がある。較正後も人間との識別能力の差が解消したとは述べていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
会話の成功度を測ることは、音声対話を人間が判定する場合でも難しい。著者らはCANDORを用い、会話の成功度を単独で採点する場合と二つを比較する場合について、先端的な大規模言語モデルを判定者として評価する。単独採点は人間の評価と中程度の相関を示す一方、対比較は長い書き起こしと提示位置による偏りの影響を受ける。判定者ごとの得点が比較しにくいため、少数の基準例と較正関数を提案し、任意の判定者の得点を共通で解釈可能な尺度へ写す。さらに、タスク指向の人間とAI、または人間とエージェントの会話200件に対比較の注釈を付けたVoice Arena Goal Datasetを公開する。このデータは、現在の判定者と人間の識別能力に大きな差があることを示す。このデータで、CANDORに合わせて作った較正が人間とAIの会話にも移るかを調べたところ、学習時にVoice Arenaのデータを見せていなくても、判定者の得点を共通尺度に乗せられた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias. Since this leaves judge scores incomparable across models, we propose a small anchor set and a calibration function that calibrates any judge onto a shared, interpretable scale. We further release the Voice Arena Goal Dataset (VA), 200 task-oriented human-AI and human-agent conversations with pairwise annotations, revealing a substantial gap between current judges and human-level discrimination. Using VA, we test whether CANDOR-fitted calibration transfers to human-AI conversations, finding it brings judges onto a shared scale despite never observing VA during fitting.
arXiv ID: 2609.29431 / 要約の誤りについて