人間評価を基準に新モデルを順位付けするANCHOR
Human-Anchored Inference for Ranking New Models with Large Language Model Judges
この論文をやさしく読む
ひとことで言うと
新しいモデルに人の評価がまだない場合、過去の人とLLM審査者の比較結果から、人基準の順位を推定する方法です。
何に役立つ?
すべての新モデルに大量の人手比較を追加する費用を抑えつつ、LLM審査者固有の偏りを補正する評価設計です。
この研究の面白いところ
審査者の感度と特徴依存の偏りを学び、それらの推定誤差の一次効果を補正します。新旧の比較で特徴分布が変わることも許容します。
どこまで分かった?
人基準の点数は推定であって、新モデルの直接の人評価ではありません。理論保証は所定の識別条件などに基づき、Chatbot Arenaでの良い誤差指標がすべての比較に自動で移るわけではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人間によるペア比較は大規模言語モデル(LLM)の評価における基準を提供するが、新しいモデルのリリースごとに十分な判定を集めるには費用と時間がかかる。LLM審判は拡張可能な代替手段だが、その比較は人間の選好や審判間で系統的に異なる可能性がある。本研究では、LLM審判による比較はあるものの人間比較はない新モデルの順位付けを調べる。 ANCHOR(ANchored Comparisons for Human-reference inference with Orthogonal Riesz correction)を提案する。これは、過去の人間比較とLLM比較から、審判ごとの人間スコア差への感度と、特徴に依存する審判バイアスを学習する。これらの推定値を新モデルの審判比較に適用し、人間を基準にしたスコアを推論する。この枠組みでは、過去の比較と新モデルの比較の間で特徴分布が変わることを許す。 推論のため、過去の人間スコア、審判感度、バイアス関数を推定することによる一次の影響を除く、共同Riesz補正を備えたNeyman直交推定量を構成する。識別、収束速度、漸近正規性を、整合的に推定できる分散とともに確立し、ANCHORが半パラメトリック効率限界に達することを示す。シミュレーションでは、スコア推定と順位付けの正確さが改善した。Chatbot Arenaでは、ANCHORは比較手法の中でスコアRMSEと挿入MAEが最も低く、平均してより狭いスコア区間を得た。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Human pairwise comparisons provide a reference for evaluating large language models (LLMs), but collecting sufficient judgments for each new release is costly and time-consuming. LLM judges offer a scalable alternative, although their comparisons may differ systematically from human preferences and across judges. We study the ranking of a new model that has received LLM-judge comparisons but no human comparisons. We propose ANCHOR (ANchored Comparisons for Human-reference inference with Orthogonal Riesz correction), which uses historical human and LLM comparisons to learn judge-specific sensitivities to human score differences and feature-dependent judge biases. These estimates are then used to infer the new model's human-reference score from its judge comparisons. The framework allows the feature distribution to change between historical and new-model comparisons. For inference, we construct a Neyman-orthogonal estimator through a joint Riesz correction that removes the first-order effects of estimating the historical human scores, judge sensitivities, and bias functions. We establish identification, convergence rates, and asymptotic normality with consistently estimable variance, and show that ANCHOR attains the semiparametric efficiency bound. Simulations demonstrate gains in score estimation and ranking accuracy. On Chatbot Arena, ANCHOR achieves the lowest score RMSE and insertion MAE among competing methods, with narrower score intervals on average.
著者のコメント
35 pages, 6 figures
arXiv ID: 2609.19599 / 要約の誤りについて