arXiv論文メモ
新着一覧
cs.LG / cs.CL · 査読状況未確認

同じ文章から作った代用ラベルによる評価を点検

Auditing Proxy-Based Validation Across Text Spans

Daein Weon, Dong Ho Kang

この論文をやさしく読む

ひとことで言うと

AIの評価スコアが安価な代用ラベルと一致しても、同じ文章の表面的な手掛かりを拾っただけかもしれないことを、文章の範囲を変えて検証した研究。

何に役立つ?

代用ラベルでスコアの妥当性を示す研究を読む際、スコアとラベルがどの文章範囲から作られたかを確かめる評価手順として役立つ。

この研究の面白いところ

HotpotQAでは50文字の接頭部分で代用ラベルとの一致が正しさとの一致を大きく上回り、OR-Benchでは冒頭の定型表現を消すと関連の大半が消えた。共通の表層情報の影響を切り分けている。

どこまで分かった?

数値は要旨に記されたHotpotQAとOR-Benchの設定に基づく。スコアと正しさの一致が偶然と同程度とされた条件も特定の短い接頭部分であり、すべての評価スコアについての結論ではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

評価スコアは、安価に得られる代用ラベルとの一致によって妥当性を確かめることが多い。しかし、スコアと代用ラベルが同じ文章の範囲から計算される場合、一致の理由は代用ラベルが表すはずの意味的な性質ではなく、両者に共通する表面的な手掛かりかもしれない。本研究では、スコア、スコアを読む文章の範囲、代用ラベル、真に測りたい性質を検証の契約として明示し、スコアを付けた範囲の外側だけで代用ラベルの規則を再評価する。共有する文章の境界だけを変えたHotpotQAの正誤判定実験では、50文字の接頭部分で、スコアと代用ラベルの一致はスコアと実際の正しさの一致を0.184上回ったが、120文字以上では差が最大0.045に縮んだ。この短い接頭部分では、スコアは後ろに回答の文字列が現れるかどうかをなお予測した(AUC 0.634)。一方、同等性検定では正しさとの一致は偶然と同程度の範囲にあり、報告された代用ラベルとの一致だけでは、スコアが正しさの順位を付けられるとは言えなかった。OR-Benchでは、各モデルで繰り返される冒頭の定型表現を除くと、スコアと拒否の代用ラベルの関連の大半が消えたが、同じ量の文章を削る対照ではほとんど消えず、測りたい性質との一致は偶然と同程度のままだった。外部の検証契約11件のうち、スコアを読む範囲外での対照を可能にするものは3件だけで、標本調査した振り分け研究はいずれも、それに必要な生成文を公開していなかった。そこで、代用ラベルによる妥当性の主張では、各ラベルを読む文章の範囲を明記し、代用ラベルとの一致と併せて本来測りたい性質との一致を報告し、スコア対象外の範囲から代用ラベルを読み直せる生成文を公開するよう求める。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Evaluation scores are often validated by their agreement with inexpensive proxy labels. When the score and the proxy are computed from the same text span, however, that agreement can arise from surface evidence the two share rather than from the semantic construct the proxy is meant to represent. We make the distinction explicit by declaring the score, its span, the proxy and the target construct as a validation contract, then re-evaluating that proxy rule strictly outside the scored span. In a controlled HotpotQA correctness experiment varying only the shared text boundary, the score agrees with its proxy far better than with correctness at a 50-character prefix: the gap is +0.184, collapsing to at most +0.045 from 120 characters onward. At that short prefix the score still predicts whether the answer string appears later (AUC 0.634) while an equivalence test places its agreement with correctness at chance, so the reported proxy agreement does not establish that the score ranks correctness. On OR-Bench, suppressing each model's recurring opening templates removes most of the score's association with the refusal proxy, while matched-volume deletion removes almost none and construct agreement stays at chance. Only three of eleven external contracts support the off-span control, and none of the routing studies we sampled released the generations it needs. We therefore ask that a proxy-based validation claim declare the span each label is read from, report the construct agreement beside the proxy agreement, and release the generations that let the proxy be re-read off the scored span.

著者のコメント

63 pages, 7 figures, 38 tables. Code: https://github.com/wdi1024/rlc-audit

arXiv ID: 2609.25808 / 要約の誤りについて