回答の提示順による偏りを補正してAIを順位付け
OSCAR: Order-aware Scoring and Calibration for AI Rankings
この論文をやさしく読む
ひとことで言うと
AI の回答を別のモデルが採点するとき、先に見せたか後に見せたかという偏りを、評価者の能力と分けて扱う手法です。
何に役立つ?
LLM の順位を比較する際、提示順や回答長などに左右される採点を補正し、比較の不確かさを見積もるために役立ちます。
この研究の面白いところ
提示順を無視すると、本来は感度のある評価者でも、感度が低いように推定され得ることを計算で示しています。また、平均だけでなく判定間の依存も補正しています。
どこまで分かった?
24.22ポイントの差は、公開表のテキスト対応付けを条件とする推定です。被覆率や RMSE の改善にはシミュレーション結果が含まれ、94.4〜95.2%は回答の正解率ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
評価者ごとの感度は、LLM の対比較評価を集約するのに有用だが、その解釈は、ランキングモデルにどの系統的な提示効果を含めるかに依存する。本研究では、AI ランキングの採点と校正のための、順序を考慮した枠組み OSCAR を導入し、そのような効果の一つとして位置を研究する。公開された18評価者の判定では、全回答にわたる A から B を引いたスコア差が、−63.11から98.31パーセントポイントに及ぶ。公開表内で質問文、回答文、候補の識別情報、評価者を対応させると、公開されたテキスト対応付けを条件として、全体の差は24.22ポイント、95%区間は[22.90, 25.54]となる。 統制した計算により、起こり得る影響を切り分ける。真の感度を1に固定したとき、値が4の位置切片を省くと、母集団で最適な傾きは0.0771に低下する。感度に基づくランキングを、評価者ごとの位置、長さ、ファミリーの項で拡張し、省略によって生じる局所的な変位と識別の失敗を特徴付け、プロンプトのクラスター単位の不確かさを補正後の比較へ伝播させる。公開された4データセットでは、位置が単独で最も大きな予測改善をもたらす。再適合を行うブートストラップ比較では、位置だけの補正に対する完全モデルの利得は、より限定的に現れる。依存のある二値シミュレーションでは、平均と共分散の両方を補正すると94.4〜95.2%の被覆率が得られ、どちらか一方の補正だけでは不十分である。N = 10,000 では、OSCAR は中立な標的に対する平均 RMSE を、感度のみのモデルの0.1158から0.0237へ低減する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effects the ranking model includes. We introduce OSCAR, an order-aware framework for scoring and calibrating AI rankings, and study position as one such effect. In released judgments from 18 evaluators, the all-response A-minus-B score difference ranges from $-63.11$ to $98.31$ percentage points. Matching question text, response texts, candidate identities, and judge within the released table gives an overall difference of $24.22$ points (95% interval $[22.90,25.54]$), conditional on the released text mapping. A controlled calculation isolates the potential consequence: with true sensitivity fixed at one, omitting a position intercept of four reduces the population-optimal slope to $0.0771$. We extend sensitivity-based ranking with judge-specific position, length, and family terms, characterize local omission-induced displacement and an identification failure, and propagate prompt-cluster uncertainty to adjusted comparisons. Across four released datasets, position provides the largest stand-alone predictive improvement. Refitting bootstrap comparisons show more selective gains from the full model over position-only adjustment. In dependent binary simulations, adjusting both the mean and covariance yields 94.4--95.2% coverage; correcting either alone is insufficient. At $N=10{,}000$, OSCAR reduces mean neutral-target RMSE from $0.1158$ under the sensitivity-only model to $0.0237$.
arXiv ID: 2609.24128 / 要約の誤りについて