LLMによるアイデアの新規性評価は指示文で揺れる
Old Ideas, Novel Problems: The Instability of LLM-Based Novelty Evaluation
この論文をやさしく読む
ひとことで言うと
LLMにアイデアの新しさを判定させると、アイデア自体が同じでも、評価時の説明の仕方で判断が大きく変わることを調べた研究です。六つの判定器で設計条件を比較しています。
何に役立つ?
研究アイデア生成システムの評価を設計・解釈する際に役立ちます。評価点が上がったときに、発想そのものの改善なのか、判定器の指示文への反応なのかを点検する材料になります。
この研究の面白いところ
同じプロンプト変更でも、判定器によって効果の向きが逆になります。専用の評価器や大きな推論予算が、安い単純なベースラインより必ず優れるわけではない点も比較しています。
どこまで分かった?
評価集合は査読者が新規性について一致した両極端の論文を選んで構築されています。境界的なアイデアや全研究分野への一般化は要旨からは判断できず、頑健な新規性評価法そのものの解決は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自動アイデア生成システムは、生成したアイデアの新規性によって評価されることが多く、その判断は大規模言語モデルに委ねられることが増えている。このような判定器は通常、その場ごとに構築され、検証されるとしても、評価すべき生成アイデアではなく人間が書いた論文で検証される。では、新規性の判定器はどの程度うまく機能するのか。十分には機能していない。 本研究では、新規性評価の設計上の選択について、系統的な統制研究を行う。まずOpenReviewから、査読者が論文の独創性を明示的に認める、または否定する文章を抽出し、研究分野内の両極端に位置する評価について全員の判断が一致した投稿だけを残すことで、評価集合を自動構築する。それらを、標準的なLLM生成器のアイデアと組み合わせる。 六つの判定器を調べると、プロンプト設計の小さな違いが大きな影響を及ぼすことが分かった。例えば、一方のアイデアは新規性があると査読者が判断し、他方はそうではないと判断した、と判定器に伝えるだけで、提示した同一のアイデア対の半数超で判定が変わり得る。対比較の正解率は50ポイント超変化し、場合によっては偶然水準を下回る。同じ変更が、ある判定器を改善し、別の判定器を悪化させる。検索や推論予算の増加はほとんど助けにならず、新規性評価のために専用設計された二つの評価器は、最も低コストなプロンプトベースのベースラインを下回った。これらの結果は、自動アイデア生成システムで報告される新規性向上に疑問を投げ掛け、頑健な新規性評価方法の必要性を示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Automated ideation systems are often evaluated on the novelty of the ideas they produce, and that judgment is increasingly delegated to large language models. Such judges are typically built ad hoc and validated, if at all, on human-authored papers rather than on the generated ideas they are meant to score. So, how do novelty judges perform? Not well. We present a systematic controlled study of novelty evaluation design choices. We first build an evaluation set automatically, mining OpenReview for passages where reviewers explicitly affirm or dispute a paper's originality and keeping only submissions with unanimous agreement at the extremes of their research area; we pair these with ideas from a vanilla LLM generator. Across six judges, we find that small prompt design choices have large consequences; e.g., simply telling the judge that reviewers found one idea novel and the other not can change its verdict on more than half of the identical idea pairs it is shown, shifting pairwise accuracy by over 50 points and occasionally pushing it below chance. The same change helps one judge and hurts another. Retrieval and larger reasoning budgets help little, and two purpose-built novelty evaluators are outperformed by our cheapest prompted baseline. These results raise questions about reported novelty gains of automated ideation systems, and call for robust novelty evaluation methods.
arXiv ID: 2610.02022 / 要約の誤りについて