AIによる文献レビューでの測定誤差と結論の信頼性
Mining Meaning: Measurement Error in AI-Assisted Literature Reviews
この論文をやさしく読む
ひとことで言うと
LLMで作った文献レビューの誤りが、論文ごとの結論と全体傾向に違う影響を与えると示した。
何に役立つ?
考えられる用途は、AI支援レビューの検証方法と、結果に基づく主張の信頼性評価である。
この研究の面白いところ
同じ誤り率でも、個別論文についての主張ほど影響を受けやすいと分析する。
どこまで分かった?
降雨を操作変数に使う経済学論文と三つの実装での事例であり、すべての分野への一般化は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
研究者は研究過程の多くの作業を自動化するため、生成AI、特に大規模言語モデル(LLM)を使うことが増えている。本研究は、学術文献を大量に読み、分類し、統合する際の信頼性を調べる。LLM支援の文献レビューを測定問題として捉え、モデルを測定系とみなして、その誤りが後の結論にどう影響するかを追う。事例として、降雨を操作変数に用いる経済学論文を特定し、メタデータを抽出するため、三種類のChatGPT実装を用いる。各実装を人がラベル付けした評価用の部分集合で検証し、全文献集合からメタデータを抽出した。LLMは二値分類では良い性能を示すが、文脈の解釈が多く必要な作業ほど性能は低下した。さらに、モデルの出力をどこまで信頼できるかは、読解課題の複雑さだけでなく、データから主張したい内容によっても変わる。同じ程度の測定誤差でも、個々の論文についての主張には大きく影響し、文献全体についての広い主張にはあまり影響しなかった。つまりLLM由来データの測定誤差が最も重要になるのは、人間のレビュアーに対するLLMの主な利点である細部の水準である。標準的なモデル性能指標は生成データの品質を知る手掛かりにはなるが、それだけで後段の推論が信頼できるとは言えない。研究者は、実質的な主張が元データを作った測定系に対して頑健かも評価する必要がある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Researchers increasingly use generative AI, particularly large language models (LLMs), to automate tasks across the research pipeline. We study the reliability of these tools at the reading, classification, and synthesis of large bodies of academic literature. We frame LLM-assisted literature reviews as a measurement problem, treating models as measurement systems and tracing how their errors affect downstream conclusions. As a test case, we use three different implementations of ChatGPT to identify and extract metadata from economics papers that use rainfall as an instrumental variable. We benchmark each implementation against a subset of human-labeled evaluation data, and then deploy those implementations to extract metadata from the full corpus. The LLMs perform well on binary classification, but performance deteriorates as tasks demand greater contextual interpretation. More importantly, how much researchers can rely on model outputs depends not only on the complexity of the reading task but also on the type of claims the data is asked to support. The same amount of measurement error substantially affects paper-level claims while having little effect on broader claims about the literature. Measurement error in LLM-generated data is thus most consequential at precisely the level of detail that constitutes an LLM's principal value added over human reviewers. We conclude that standard model performance metrics are informative about the quality of generated data but do not by themselves establish the credibility of downstream inference. Researchers must also evaluate whether substantive claims are robust to the measurement system used to generate the underlying data.
arXiv ID: 2609.27686 / 要約の誤りについて