arXiv論文メモ
新着一覧
stat.ML / cs.LG / stat.AP · 査読状況未確認

回答の提示順による偏りを補正してAIを順位付け

OSCAR: Order-aware Scoring and Calibration for AI Rankings

You Liu, Yue Liu, Quanchao Lu, Nick Shipilov

この論文をやさしく読む

ひとことで言うと

AI の回答を別のモデルが採点するとき、先に見せたか後に見せたかという偏りを、評価者の能力と分けて扱う手法です。

何に役立つ?

LLM の順位を比較する際、提示順や回答長などに左右される採点を補正し、比較の不確かさを見積もるために役立ちます。

この研究の面白いところ

提示順を無視すると、本来は感度のある評価者でも、感度が低いように推定され得ることを計算で示しています。また、平均だけでなく判定間の依存も補正しています。

どこまで分かった?

24.22ポイントの差は、公開表のテキスト対応付けを条件とする推定です。被覆率や RMSE の改善にはシミュレーション結果が含まれ、94.4〜95.2%は回答の正解率ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

評価者ごとの感度は、LLM の対比較評価を集約するのに有用だが、その解釈は、ランキングモデルにどの系統的な提示効果を含めるかに依存する。本研究では、AI ランキングの採点と校正のための、順序を考慮した枠組み OSCAR を導入し、そのような効果の一つとして位置を研究する。公開された18評価者の判定では、全回答にわたる A から B を引いたスコア差が、−63.11から98.31パーセントポイントに及ぶ。公開表内で質問文、回答文、候補の識別情報、評価者を対応させると、公開されたテキスト対応付けを条件として、全体の差は24.22ポイント、95%区間は[22.90, 25.54]となる。 統制した計算により、起こり得る影響を切り分ける。真の感度を1に固定したとき、値が4の位置切片を省くと、母集団で最適な傾きは0.0771に低下する。感度に基づくランキングを、評価者ごとの位置、長さ、ファミリーの項で拡張し、省略によって生じる局所的な変位と識別の失敗を特徴付け、プロンプトのクラスター単位の不確かさを補正後の比較へ伝播させる。公開された4データセットでは、位置が単独で最も大きな予測改善をもたらす。再適合を行うブートストラップ比較では、位置だけの補正に対する完全モデルの利得は、より限定的に現れる。依存のある二値シミュレーションでは、平均と共分散の両方を補正すると94.4〜95.2%の被覆率が得られ、どちらか一方の補正だけでは不十分である。N = 10,000 では、OSCAR は中立な標的に対する平均 RMSE を、感度のみのモデルの0.1158から0.0237へ低減する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Judge-specific sensitivity is useful for aggregating pairwise LLM evaluations, but its interpretation depends on which systematic presentation effects the ranking model includes. We introduce OSCAR, an order-aware framework for scoring and calibrating AI rankings, and study position as one such effect. In released judgments from 18 evaluators, the all-response A-minus-B score difference ranges from $-63.11$ to $98.31$ percentage points. Matching question text, response texts, candidate identities, and judge within the released table gives an overall difference of $24.22$ points (95% interval $[22.90,25.54]$), conditional on the released text mapping. A controlled calculation isolates the potential consequence: with true sensitivity fixed at one, omitting a position intercept of four reduces the population-optimal slope to $0.0771$. We extend sensitivity-based ranking with judge-specific position, length, and family terms, characterize local omission-induced displacement and an identification failure, and propagate prompt-cluster uncertainty to adjusted comparisons. Across four released datasets, position provides the largest stand-alone predictive improvement. Refitting bootstrap comparisons show more selective gains from the full model over position-only adjustment. In dependent binary simulations, adjusting both the mean and covariance yields 94.4--95.2% coverage; correcting either alone is insufficient. At $N=10{,}000$, OSCAR reduces mean neutral-target RMSE from $0.1158$ under the sensitivity-only model to $0.0237$.

arXiv ID: 2609.24128 / 要約の誤りについて