arXiv論文メモ
新着一覧
cs.SI · 査読状況未確認

コミュニティノートの意見対立は一本の軸で足りるか

Two Fault Lines: Latent Polarity Geometry in X Community Notes

Andreas Andreou, Michael Sirivianos

この論文をやさしく読む

ひとことで言うと

Xのコミュニティノートで意見の違いを政治的な左右だけで捉えると、別の対立が見えなくなるという分析です。公開評価データから、制度への信頼と解釈される第二の軸を見いだしています。

何に役立つ?

ノートの評価モデルを改良したり、評価者が不足する言語を把握したりする参考になります。複数軸への変更や評価者募集は著者の提案で、導入効果を実験した結果ではありません。

この研究の面白いところ

第二の軸は、学習から外したCOVIDやウクライナのノートへの判断にも関係しました。言語ごとの差についても、評価数の不足と公開ルールの違いを分けて調べています。

どこまで分かった?

公開評価データのモデル再適合と予測評価に基づく分析です。第二の軸を制度への信頼と呼ぶのは著者の解釈であり、相関や公開率の差をそのまま因果効果とは扱えません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

コミュニティノートは、Xの利用者参加型ファクトチェックシステムである。普段は意見が合わない評価者の双方が役に立つと評価した場合にのみ、訂正対象の投稿の下にノートが表示される。この設計はブリッジングと呼ばれる。このルールを適用するため、システムは評価だけから誰と誰の意見が対立するかを学習し、すべての評価者とノートを、極性軸と呼ばれる一本の直線上に配置する。本番の処理系にあるすべてのスコアラーが単一の軸を用いている。 これらのスコアラーが共有する基礎モデルを全公開データ、すなわち2億1290万件の評価、233万件のノート、107万人の評価者に再適合したところ、軸は一本では足りないことが分かった。この空間は少なくとも二次元である。第一の軸は政治的な左派・右派であり、第二の軸は、私たちが制度への信頼と解釈するもので、第一の軸からおおむね独立している。学習に使わないデータでのテストは、第二の軸が未知の評価の予測を改善する一方、第三の軸による追加効果は小さいことを確認した。ある一組の話題から学習した評価者の第二の次元は、適合から除外したCOVIDとウクライナのノートを評価者がどう判断するかを予測するため、単に話題の違いを言い換えたものではない。 評価数が多く、政治的には評価者をほとんど分断しないノートの中で、第二の軸に沿った不一致が大きくなると、公開される割合は71.5%から11.7%へ低下する。単一軸の適合では、これらのノートは極性が弱く有用性も低いとしか記録されず、第二の軸の一方の端にいる評価者が支持しているという情報が失われる。執筆者は両方の軸で自分の立場に合うノートを書いており、相関係数rはそれぞれ0.538と0.358である。また、少数の評価者が評価の大部分を行っている(ジニ係数0.718)。 最も小規模な言語コミュニティでは公開されるノートが少ないが、不足しているのは受け取る評価数であり、ルールの扱い方ではない。ノート当たりの評価数を一定にすると、全体の公開率10.85%を下回るのはヒンディー語だけとなり、ギリシャ語は7.76%から11.68%へ上がる。私たちは、意見の不一致を表す軸を複数持つブリッジングモデルと、現在の設計が最も届きにくい言語での評価者の募集を提唱する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Community Notes is X's crowdsourced fact-checking system. A note is published beneath the post it corrects only when raters who usually disagree both rate it helpful, a design called bridging. To apply that rule, the system learns who disagrees with whom from the ratings alone, placing every rater and note on one line, the polarity axis. Every scorer in the production pipeline uses a single axis. Refitting the base model these scorers share on the full public data (212.9M ratings, 2.33M notes, 1.07M raters), we find that one axis is too few. The space is at least two-dimensional. The first axis is left/right politics, while the second, which we interpret as trust in institutions, is largely independent of the first. A held-out test confirms that the second axis improves prediction of unseen ratings, while a third adds little. A second rater dimension learned from one set of topics predicts how raters judge COVID and Ukraine notes excluded from the fit, so it does not merely restate subject matter. Among heavily rated notes that barely divide raters politically, the published share falls from 71.5% to 11.7% as second-axis disagreement grows. A one-axis fit records these notes only as weakly polarised and less helpful; the information that raters at one end of the second axis support them is lost. Authors write notes matching their own position on both axes (r = 0.538 and 0.358), and a small minority of raters cast most ratings (Gini = 0.718). Fewer notes are published in the smallest language communities, but the shortfall is in ratings received, not in how the rule treats them. Keeping ratings per note constant, only Hindi stays below the global rate of 10.85%, and Greek moves from 7.76% to 11.68%. We argue for a bridging model with more than one axis of disagreement, and for recruiting raters in the languages the current design reaches least.

arXiv ID: 2609.21496 / 要約の誤りについて