arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

長文の価値観判定で文の数より証拠の強さを集約

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

Yuhe Wu, Rui Qian, Guangyu Wang, Yuran Chen, Yuanchao Zhu, Junjie Yang, Zhengheng Li, Jiulin Cai, Tianyi Zhang, Zihan Dong, Jiaxin Liu, Yujie Chen, Guang Zhang

この論文をやさしく読む

ひとことで言うと

長い投稿から価値観を測るとき、文の多数決ではなく各文が持つ証拠の強さで判定をまとめる方法を示した。

何に役立つ?

引用や背景説明が多い長文を言語モデルで分類するとき、不確かな文に引きずられにくい集約法として検討できる。

この研究の面白いところ

中国語・英語8,358投稿のベンチマークで、比較法より平均4.5ポイント高い正解率と4.6ポイント高いマクロF1を示した。

どこまで分かった?

結果はMINDの六つの価値次元、五つのモデル、二つの言語での評価であり、あらゆる長文分類への一般化は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルは、長いソーシャルメディア投稿から人々の価値観の方向を測るために使われることが増えている。しかし、こうした投稿には背景説明、引用、譲歩が混在し、立場を示す文は少数だけの場合が多い。既存の方法は、文書全体のラベルを直接予測して過信するか、文ごとの予測を多数決や確率の平均で集約して、不確かな文と決定的な文を同じように扱う。本研究は、長文の価値観測定を判断の融合問題として定式化し、学習を必要としないTempered Evidence Fusion(TEF)を提案する。TEFは、一般化ベイズ事後分布から導いた、各文の正規化された情報利得に応じて、その文の対数オッズを重み付けする。そのため、不確かな文の寄与はほぼ消え、決定的な証拠にはベイズ最適な重みが残る。さらに、過去5年間の公共的な出来事に関する中国語と英語の投稿8,358件、六つの価値次元を含むベンチマークMINDを導入する。MINDでは、五つの言語モデルと二つの言語にわたり、TEFは直接判定、多数決、確率平均のうち最も強い基準法より、平均で正解率が4.5ポイント、マクロF1が4.6ポイント高かった。MINDのデータとコードは公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.

著者のコメント

Yuhe Wu, Rui Qian, and Guangyu Wang contributed equally. Corresponding author: Guang Zhang. See also: https://github.com/Kzczc/ICASSP2027-TEF

arXiv ID: 2609.27165 / 要約の誤りについて