arXiv論文メモ
新着一覧
cs.CL / cs.CY · 査読状況未確認

社会科学の文章ラベル付けで判断専用モデルを比較

Evaluating Decision Models for Text Annotation in Computational Social Science

Hazem Ibrahim and Yasir Zaki

この論文をやさしく読む

ひとことで言うと

社会科学の文章分類で、回答と確率を返す判断専用モデルの精度・費用・自信度を比較した研究です。

何に役立つ?

大量の文章をラベル付けする際、低自信度の項目だけ高価なモデルに回す設計の参考になります。

この研究の面白いところ

判断モデル単独では多くの課題で先端モデルに劣りますが、費用は大幅に低く、振り分けて組み合わせると費用を抑えられました。

どこまで分かった?

一部の課題では高自信度でも偶然水準に近い結果でした。自信度だけで正しさを保証できず、課題ごとの検証が必要です。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

計算社会科学では文章へのラベル付けに大規模言語モデルを使うことが増え、発表された結論の妥当性はそのラベルに依存する。新しい種類の判断モデルは、自由文ではなく、型が決まった質問に対して選択肢、ラベル全体の確率分布、自信度を返し、先端モデルの推論料金のごく一部で使える。しかし社会科学の概念について、その答えが正確か、申告された自信度を信用できるかは分かっていない。本研究はZiemsらの2024年の評価をなぞり、計算社会科学の分類18課題、7,977項目について、最初の商用判断モデルと二つの公開重みモデルを、19の先端および公開重みの言語モデルと同じゼロショット手順で比較した。さらに、その後一週間に公開された11の公開重みシステムにも判断モデルの比較を広げた。判断モデルは評価可能な15課題のうち14課題で、課題ごとに最良の言語モデルより劣り、Macro-F1の差の中央値は11.6ポイントだったが、測定した費用の中央値は44分の1だった。自信度の較正は19の言語モデルのうち16モデルの言語化された自信度より良かったものの、三つの先端モデルでは較正誤差の中央値が判断モデルの0.157より低い0.066だった。自信度0.9を超える項目は通常は正確にラベル付けされ、正解率中央値は0.815だった。ただし、仲間同士の支援対話における共感を扱う一課題では、高い自信度を示しながら成績は偶然水準に近かった。それでも、判断モデルはラベル付けの最初の段階で有用と示唆される。自信度が低い項目だけを言語モデルへ回すと、費用を4分の1から2分の1に抑えながら、言語モデル単独と同等以上の結果になった。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the labels generated by such models. Decision models, a new model class built for categorical question answering, answer typed questions with a choice, a probability distribution over the label set, and a confidence score rather than free text, at a small fraction of frontier inference prices. Whether their answers are accurate, and whether that stated confidence can be trusted on social science constructs, are unknown. Here, we mirror the evaluation of Ziems et al. (2024) on 18 computational social science classification tasks (7,977 items), comparing the first commercial decision model and two open-weight counterparts against 19 frontier and open-weight language models under the same zero-shot protocol, and extending the decision-model comparison to eleven open-weight systems released in the week after it. The decision model trails the per-task best LLM on 14 of 15 evaluation tasks, with a median deficit of 11.6 macro-F1 points, at a median 44 times lower measured cost. Its confidence is better calibrated than the verbalized confidence of 16 of the 19 LLMs, yet three frontier models show lower median calibration error (0.157 against 0.066). While items above 0.9 confidence are typically labeled accurately (median accuracy 0.815), on one task, empathy in peer-support dialogues, the model reports high confidence while performing near chance. Nonetheless, our results suggest that decision models are useful as a first step in the annotation pipeline: routing low-confidence items to an LLM matches or exceeds the LLM alone at a quarter to half of its cost.

著者のコメント

54 pages, 8 figures, 25 tables

arXiv ID: 2609.24574 / 要約の誤りについて