査読文に含まれる外部情報を測る自己条件付け法
How Much Were You Told? Measuring External Information in Peer Reviews
この論文をやさしく読む
ひとことで言うと
査読文がAIらしい書き方かどうかではなく、元論文や一般的な指示だけでは説明できない情報をどれほど含むかを測った。
何に役立つ?
AIによる文章の推敲と批評内容の委任を区別するための評価法として考えられる。学会の実際の判定制度での有効性を実証したとは要旨にない。
この研究の面白いところ
査読文自身から抽出したヒントを生成文脈に足す前後で、文のもっともらしさを比べる。表面の言い換えに比較的強い点が従来法と違う。
どこまで分かった?
AUC最大1.0はIntelLabsの査読ベンチマークでの結果である。高温度サンプリングによる回避が可能で、品質低下を伴うと報告されている。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
学会の方針は、大規模言語モデル(LLM)を自分の査読文の推敲に使うことと、批評自体を委ねることを区別している。しかし、現在の人工生成文検出法(ATD)は、内容の出所よりも表面的な書き方を主に測っている。本研究は代わりに、査読対象の論文と一般的な査読指示だけでは説明できない、査読文に含まれる外部情報を測る。 自己条件付け(Self-Conditioning)という教師なしの情報理論的推定法を提案する。査読文が、その生成時の文脈の下でどの程度もっともらしいかを、査読文自体から取り出したヒントを文脈に追加した場合のもっともらしさと比較する。IntelLabsの査読ベンチマークでは、批評全体を委ねた査読文と機械的に推敲した査読文をAUC最大1.0で区別し、表面的な言い換えにはおおむね影響されなかった。生成器に外部から与える情報を増やすと、この方法の得点は人間の査読文の領域へ単調に近づき、標準的なATDの比較手法とは異なる挙動を示した。一方、高温度でのサンプリングによって推定法を回避できるが、その際には出力品質が下がる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Conference policies distinguish using Large Language Models (LLMs) to polish one's own review from delegating the critique, but current Artificial Text Detection (ATD) methods largely measure surface form rather than the origin of its content. We instead measure the external information carried by a review: information not explained by the reviewed paper and a generic reviewing instruction. We propose Self-Conditioning, an unsupervised information-theoretic estimator that compares the likelihood of a review under its production context with its likelihood when that context is augmented with hints extracted from the review itself. On the IntelLabs peer-review benchmark, Self-Conditioning separates fully-delegated from machine-polished reviews with AUC up to $1.0$ while remaining largely insensitive to surface rewriting. Moreover, as generators receive increasing amounts of externally-provided information, their scores move monotonically towards the human regime, unlike standard ATD baselines. High-temperature sampling can evade the estimator, but at the cost of output quality.
arXiv ID: 2609.28041 / 要約の誤りについて