arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

政策文書から調査回答を作る言語モデルの評価

From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring

Carolyn Cole, Matthias Deschryvere, Toqeer Ehsan, Arash Hajikhani

この論文をやさしく読む

ひとことで言うと

国ごとの政策文書を言語モデルで読み取り、政策手段や対象などの調査項目へ整理する方法を評価します。

何に役立つ?

多くの国の政策を比較・追跡する際の情報整理に役立つ可能性があります。人による確認と文脈解釈は必要とされています。

この研究の面白いところ

構造化された項目では人の回答と84~95%一致しましたが、自由記述ではモデルが手続きの詳細を多く書く違いがありました。

どこまで分かった?

評価は複数国のデータセットでの人との一致に関するものです。自由記述の差が残り、完全自動化を裏付ける結果ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

科学、技術、イノベーションの政策は競争力に重要だが、政策の種類が多く規模も大きいため、一貫して整理し追跡するのは難しい。既存の方法は人手の調査に大きく依存し、費用がかかるうえ、国をまたいだ規模拡大も難しい。大規模言語モデル(LLM)は、長く構造化されていない政策文書から情報を抽出して整理する新たな可能性を与える。本論文は、政策文書から構造化された調査回答を作る「AI回答者」としてLLMを使う方法を提示する。 長い文脈を扱う文脈内学習に基づくデータ抽出処理を開発し、公開ウェブ情報を政策手段、対象集団、テーマ領域など、定められた調査分類に対応付ける。この処理には、関連性と根拠を評価する別のLLMによる検証段階と、人が作成した回答との比較も組み込む。複数国のデータセットを使い、回答の重なりを測る指標と交差検証によって、LLMと人による出力の一致を評価した。構造化された指標では84~95%と高い一致が得られたが、自由記述欄には違いが残り、モデルは手順についてより詳細に書く傾向があった。この結果は、人による検証と文脈の解釈を引き続き必要としながら、人とAIを組み合わせた作業によって政策監視の効率と規模を改善できる可能性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDFDOI

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as "AI respondents" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.

著者のコメント

Accepted as a full paper to FLINS-ISKE 2026

arXiv ID: 2609.29370 / 要約の誤りについて