arXiv論文メモ
新着一覧
cs.AI / cs.CL / cs.CY / q-fin.RM · 査読状況未確認

LLM評価の再現性は測りたい概念の妥当性を保証しない

Reproducibility is not construct validity: LLM measurement of institutionally situated communication

Veronika Batzdorfer (KIT), Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, médialab, Sciences Po)

この論文をやさしく読む

ひとことで言うと

LLMの採点が何度やっても同じでも、本当に測りたい態度を測れているとは限らないことを、AI法の意見募集データで検証しています。

何に役立つ?

自由記述から人や組織の態度を数値化する研究で、再現性と測定の妥当性を別々に確認する必要性を示します。

この研究の面白いところ

同じ利害関係者のアンケートと意見書を対応させ、LLM注釈の級内相関が0.99超でも両者の一致は限定的でした。業界団体は自由記述の方がAIリスクへの懸念を強く示すなど、差は集団によって異なりました。

どこまで分かった?

欧州委員会の一つの意見募集に基づく関連の分析です。どちらの回答形式が真の態度を完全に表すかを確定したわけではなく、制度的な発言文脈も区別する必要があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

アノテーションの再現性が高くても、LLMが推定した尺度が、測定を意図した概念を捉えているとは限らない。本研究では、欧州委員会によるAI法の意見募集のデータセットを用い、同じ利害関係者による構造化された調査回答と自由記述の意見書を結び付けて、この違いを検証する。意見書に対するLLMのアノテーションは極めて高い再現性を示す(級内相関係数は0.99超)。しかし、本来近似しようとした概念について、調査回答から得た尺度との収束は限定的である。 調査に基づく尺度とLLMが文章から推定した尺度の乖離は、利害関係者の集団ごとに系統的に異なる。業界団体は、調査回答よりも文章による意見募集でAIのリスクへの強い懸念を示す(ḡ = +1.0)。一方、公的機関や複数の非企業集団では、乖離は小さいか負である。スコア間の乖離には、欧州諸国間で正の空間的自己相関が示唆される(MoranのI = 0.347、p = 0.036)。これは、隣接国の利害関係者ほど、AIの安全性への懸念に関する文章上の立場が似る傾向を示している。乖離があっても、調査で報告された懸念は、どの乖離水準でも説明可能性への支持と強く関連している。 以上の結果は、LLMによるアノテーションの再現性と、測定概念との対応の乏しさが両立し得ることを示す。LLMを測定手段として用いる際には、再現性、構成概念妥当性、コミュニケーションの文脈による変動を区別する検証手順が必要である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses (ḡ = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.

arXiv ID: 2609.19866 / 要約の誤りについて