記事の判定を集約して報道媒体の信頼性を推定する
From Articles to Publishers: Aggregating Language Model Predictions for News Source Reliability Inference
この論文をやさしく読む
ひとことで言うと
一つの記事だけで媒体全体を評価せず、複数記事に対する言語モデルの判定をまとめて媒体の信頼性を推定します。学習に使った媒体と評価対象の媒体を分けて試しています。
何に役立つ?
考えられる用途は、多数の媒体を調査する際に文章面から評価候補を整理することです。専門機関の評価を教師として推定する研究であり、記事ごとの事実確認を代行できると実証したものではありません。
この研究の面白いところ
評価対象を記事から媒体に変えることで、個別記事の判定の揺れを集約する設計です。精度だけでなく、政治的傾向と誤り方との関連も分析しています。
どこまで分かった?
対象は英語の政治記事で、正解ラベルはNewsGuardの評価です。0.60と0.69は記事単位と媒体単位という異なる評価単位の正解率です。政治的立場による誤分類の違いがあり、無偏りな判定を確立したわけではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
報道媒体の信頼性は従来、編集方針、透明性、事実に関する基準を発信元で評価する専門機関によって判定されてきた。この過程を計算的な手法に置き換える際には、個々の記事を対象とする問題として定式化し、あらかじめラベルを付けた記事群でモデルを学習して、テスト段階で性能を評価することが多い。 本研究では、ニュース発信元の信頼性推定を、発信元単位の予測問題として調べる。Transformerに基づく言語モデルがまず個々の記事の信頼性を推定し、続いて記事単位の予測を集約して、学習で見ていない媒体の信頼性を推定する2段階の枠組みを提案する。現実的な運用条件に近づけるため、学習集合とテスト集合に同じ媒体が一切含まれない、媒体を厳密に分離した評価手順を採用する。 NewsGuardの信頼性評価をラベルとした、英語媒体439社の政治ニュース19,476記事を用いた実験では、集約により頑健性と性能が大きく改善し、正解率は記事単位の約0.60から媒体単位の0.69へ上昇した。最後に、政治的立場によって予測誤りがどう変わるかを分析し、政治的傾向と誤分類パターンの間に統計的に有意な関連を見いだした。全体として、これらの結果は、文章から得られる信号を集約するだけで媒体の信頼性を推定できることを示し、規模を拡張できる内容ベースのニュース発信元自動評価を支持する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Traditionally, the reliability of news publishers is assessed by expert organisations that evaluate editorial practices, transparency and factual standards at source. When this process is translated into a computational approach, the problem is often formulated at the level of individual articles, with models being trained on a set of pre-labelled articles and their performance being evaluated in a test phase. In this work, we investigate news source reliability inference as a source-level prediction problem. We propose a two-stage framework in which transformer-based language models first estimate the reliability of individual articles and subsequently aggregate article-level predictions to infer the reliability of previously unseen publishers. To approximate realistic deployment conditions, we enforce a strict publisher-disjoint evaluation protocol, ensuring that no publisher appears in both training and test sets. Experiments on 19,476 political news articles from 439 English-language publishers labeled with NewsGuard reliability ratings show that aggregation substantially improves robustness and performance, increasing accuracy from approximately 0.60 at the article level to 0.69 at the publisher level. Finally, we analyze how prediction errors vary across political orientations, revealing statistically significant associations between political leaning and misclassification patterns. Overall, our findings show that publisher reliability can be inferred from aggregated textual signals alone, supporting scalable and content-based approaches to automated news source assessment.
著者のコメント
10 pages, 7 figures, 2 tables. Submitted to IEEE Transactions on Computational Social Systems (TCSS)
arXiv ID: 2609.24219 / 要約の誤りについて