道徳的判断の不一致を残してラベル集約の偏りを調べる
Moral Entropy: Auditing Bias and Uncertainty in Moral Judgment
この論文をやさしく読む
ひとことで言うと
道徳的な内容かどうかで人の判断が割れたとき、その割れ方を消さずに保持し、多数決などのまとめ方が生む偏りを調べます。
何に役立つ?
倫理関連データのラベルを作る際に、拾いすぎと見逃しを評価する材料になります。不一致自体を不確実性として扱える点が重要です。
この研究の面白いところ
一人でも肯定すれば採用する規則と、複数票を求める規則で、誤りの方向が大きく異なります。全体集計と道徳的基盤別の集計でも見え方が変わります。
どこまで分かった?
比較基準は枠組みが推定した較正済み事後分布です。数値を道徳上の絶対的な真偽との誤差とみなすことはできません。結果は三コーパス・15分野について報告されています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
計算倫理の研究の多くは、道徳的内容についての注釈者間の不一致を、投票で取り除く雑音として扱う。多数決で集約したり、一人でも該当すると判断した時点で採用する、より緩いany-annotator規則を用いたりする。本研究では、この不確実性をモデル化して学習に利用すべきだと論じる。 提案するMoral Entropyは、真のラベルに関する事後分布全体を保持し、そのエントロピーを、偶然的不確実性、すなわち道徳的内容に関する取り除けない不一致と、認識的不確実性、すなわち注釈の不足や雑音に由来する不確実性に分解するベイズ的枠組みである。また、交差エントロピーやKL、Brierスコア、期待較正誤差などのエントロピーに関わる評価手法を使い、任意の経験的な合意規則を較正された正解基準と照らして監査できる。 三つのコーパスと15の言説分野で、標準的な集約規則をこの事後分布と比較すると、現在の処理手順では報告されていない偏りが明らかになる。any-annotator規則は約30%の項目で較正済み事後分布と食い違い、全体をまとめるとほぼすべてが偽陽性である。ただし道徳的基盤別に見ると誤りの傾向は逆転し、MFTCでは平均偽陽性率19.9%、平均偽陰性率38.9%となる。一方、より厳しい多数決規則と2票規則は、真の陽性の63〜83%を見逃す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Most work in computational ethics treats annotator disagreement on moral content as noise to be voted away, collapsed into majority vote or the more permissive any-annotator rule the moment a single annotator flags an item. We argue this uncertainty should instead be modeled and learned from. We introduce Moral Entropy, a Bayesian framework that keeps a full posterior over the true label and decomposes its entropy into aleatoric uncertainty (irreducible disagreement about the moral content) and epistemic uncertainty (from insufficient or noisy annotation) -- and lets any heuristic consensus rule be audited against a calibrated ground truth via entropy methods such as cross-entropy/KL, Brier score, and expected calibration error. Across three corpora and fifteen discourse domains, auditing the standard aggregation rules against this posterior reveals bias that no current pipeline reports: the any-annotator rule disagrees with the calibrated posterior on roughly 30% of items -- pooled, almost entirely false positives, though the errors invert at the foundation level (19.9%/38.9% mean FPR/FNR on MFTC) -- while the stricter majority and two-vote rules miss 63-83% of true positives.
著者のコメント
accepted to UncertaiNLP @ EMNLP 2026
arXiv ID: 2609.21992 / 要約の誤りについて