arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

AIによる査読評価指標は文章の見た目に影響されるか

Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics

Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri

この論文をやさしく読む

ひとことで言うと

AIで査読文の表現だけを整えると、内容が同じでも評価点が変わるか調べた。

何に役立つ?

人間とAI支援の査読を比較する評価指標を選ぶ際、文章表現への偏りを確認する助けになる。

この研究の面白いところ

674件の元査読から4,044件の書き換えを作り、29指標中23指標が同内容でも有意に点数を変えた。

どこまで分かった?

意味を保つ書き換えという比較条件での結果であり、査読の真の品質そのものを直接測ったとは限らない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

査読の評価に、大規模言語モデルを判定者とする指標が増えているが、そこには測定上のリスクがある。論文の評価内容が優れているからではなく、文章が流暢で整理され、洗練されているために高得点を得るかもしれない。特に、査読者が評価内容を変えずに文章だけをAIで改善する場合に問題となる。本研究は、査読評価指標が表面的な言語表現を超えて実質的な査読品質を捉えるか検査する、統計的な枠組みを提案する。同じ評価内容を保ちながら表現と提示方法を変えた、元の人間による査読とLLMによる書き換えを比較する。 ICLRとNeurIPSの人間による査読674件から作った、意味を保つ書き換え4,044件のデータセットを使い、先行研究4件から集めた内容重視の査読評価指標29種類を、表現への感度と頑健性を補完的に調べる検査で評価した。これらの指標は表面的な文章特性を超える査読の性質を捉えることを意図しているが、書き換えへの感度は広く見られた。主要な解析では、評価内容を保った査読に23指標が有意に異なる得点を与え、頑健性基準を満たしたのは6指標だけだった。傾向は二つのLLM判定者でおおむね一致し、一つの判定者に固有の問題ではないことを示唆する。多くの指標が査読内容の質と文章表現を部分的に混同しており、人間執筆、AI支援、AI生成の査読を比較する前に、意味を保つ書き換えへの頑健性を検証すべきだと結論づける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted reviewing, where reviewers may use LLMs to improve clarity or presentation while preserving the underlying judgments. We propose a statistical framework for testing whether peer-review evaluation metrics capture substantive review quality beyond surface-level linguistic form. The framework compares original human reviews with faithful LLM rewrites that preserve the same evaluative content while changing wording and presentation. Using a dataset comprising 4,044 meaning-preserving rewrites derived from 674 human reviews from ICLR and NeurIPS, we evaluate 29 content-oriented peer-review evaluation metrics drawn from four prior works through complementary tests of surface sensitivity and robustness. Although these metrics are intended to capture review properties beyond surface-level, writing-dependent characteristics, we find that sensitivity to rewriting is widespread. Under our primary analysis, 23 metrics assign significantly different scores to reviews whose evaluative content is preserved, while only six satisfy our robustness criterion. The patterns are largely consistent across two LLM judge models, suggesting that the issue is not specific to a single judge. These findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.

著者のコメント

Accepted at CIKM 2026

arXiv ID: 2609.23264 / 要約の誤りについて