研究助成申請の採点に言語モデルを使えるかを比較
Scoring Grant Applications with Large Language Models
この論文をやさしく読む
ひとことで言うと
英国の研究助成申請2267件を六つの言語モデルで採点し、人間の査読者や審査委員の評価と比べた。
何に役立つ?
助成審査の初期段階でAIの採点が補助として使えるかを考えるための実測データになる。
この研究の面白いところ
最良のモデルでも査読者平均との順位相関は0.26で、審査委員の平均との相関は0.17にとどまった。
どこまで分かった?
個々のモデル採点は不正確で、最終審査で専門家の判断を代替するには相関が弱い。初期審査への利用は可能性として挙げられている。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
研究助成の申請書を評価する作業は時間がかかり、学術的な査読全体の負担を増やす。助成機関はAIが助けになるかを検討しているが、最近の助成申請を大規模言語モデル(LLM)がどれほど正確に採点できるかについて、公表された研究はなかった。本研究は、公開重みの六つのモデル(Gemma 3の1B、4B、12B、27B、DeepSeek R1 32B、Qwen 3 32B)で、英国の経済社会研究会議(ESRC)と工学・物理科学研究会議(EPSRC)の最近の申請2267件を採点し、元の査読者や助成審査委員の点数と比較した。 モデルの点数は個別には不正確だったが、平均して順位に変換すると専門家の平均点と正の相関を示した。最も成績が良かったGemma 3 27Bを、プロンプトを変えて10回実行すると、査読者の平均点との順位相関は中程度で、平均ρ=0.26だった。個々の査読者との平均相関は0.19で、査読者同士の平均相関0.24より低く、モデルの採点は個々の査読者よりやや弱いことが示唆された。審査委員の平均点との順位相関は弱く、平均ρ=0.17で、個々の委員との平均相関はρ=0.14だった。これは委員同士の相関ρ=0.38を大きく下回る。最終的な委員会審査で専門家を置き換えるには相関が弱すぎるとみられる一方、弱い申請の早期却下の候補を探す、査読者一人分の代わりにする、偏りを確認するために別の視点を加えるなど、初期審査では役立つ可能性がある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Purpose: Assessing grant applications is time-consuming and difficult, adding to the overall burden of academic peer review. Whilst funders are exploring whether AI can help, there is no published research into the accuracy of Large Language Models (LLMs) for scoring contemporary grants. Design/methodology/approach: This study investigates whether six open-weight LLMs (Gemma 3 1B/4B/12B/27B, DeepSeek R1 32B, Qwen 3 32B) can give useful scores for 2267 recent UK Economic and Social Research Council (ESRC), and Engineering and Physical Sciences Research Council (EPSRC) grant applications, comparing them with scores from the original reviewers and funding panel members. Findings: Although the LLM scores are individually inaccurate, when averaged and converted to ranks they correlate positively with expert average scores. The best performing LLM, Gemma 3 27B (10 iterations with varied prompts), had moderate rank correlations with average reviewer scores (mean rho=0.26). Gemma 3 27B's average correlation with individual reviewers was 0.19, which is lower than the inter-reviewer mean correlation of 0.24, suggesting that it scores are slightly weaker than individual reviewer scores. Gemma 3 27B had weak rank correlations with average panel member scores (mean rho=0.17), with lower average correlations with individual panellists (mean rho=0.14), which is substantially lower than the inter-panellist correlation (mean rho=0.38). Whilst the correlations seem too weak to replace expert review at the final panel stage, LLM scores might help with the initial reviewing state, such as by helping identify the weakest proposals for fast-track desk rejections, to replace one human reviewer, or for triangulation to check for bias.
arXiv ID: 2609.25327 / 要約の誤りについて