翻訳の自動評価は職業と性別で偏るか
Benchmarking Gender Bias in Machine Translation Evaluation Metrics across Occupations
この論文をやさしく読む
ひとことで言うと
原文に性別が書かれていないのに、男性形と女性形の訳で自動評価の点が違うかを調べています。職業を均等に含むデータを使い、7つの翻訳先言語で比較します。
何に役立つ?
翻訳システムを採点する側の偏りを点検するのに役立ちます。平均点だけでなく、偏りの方向や頻度、職業別の傾向も確認する必要性を示しています。
この研究の面白いところ
翻訳そのものの偏りではなく、その翻訳を評価する指標の選好に焦点を当てています。436の職業群に同数の例を用意し、職業構成の影響を抑えています。
どこまで分かった?
対象は英語原文と7言語への翻訳、共有タスクの提出システムおよびベースラインです。男性形を高く評価する傾向は一様ではなく、評価器と言語で強さや一貫性が異なります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ジェンダーバイアスは機械翻訳(MT)で引き続き懸念されており、生成された翻訳とその自動評価の双方に影響する。原文が人物の性別を特定しない場合、翻訳は男性形または女性形でその人物を表すことがあり、原文にはその区別の根拠がなくても、MTシステムと評価指標の双方がこれらの選択肢に系統的な選好を示し得る。 本研究では、職業のバランスを取ったGAMBIT+の部分集合を用い、WMT 2026の自動翻訳品質評価システム共有タスクで、この挙動を調べる。英語を原言語とする7言語対を対象とし、元のデータセットにあるアラビア語、チェコ語、ギリシャ語、アイスランド語、ロシア語、ウクライナ語の6つに、ドイツ語を追加して元の資料を拡張する。この部分集合には、翻訳先の言語ごとに男性形・女性形の翻訳対が1,308組含まれ、ISCO-08の436職業群それぞれに3例を用意する。 共有タスクへの提出システムとベースラインについて、スコア予測および誤り注釈を評価し、性別に関係する差の方向、大きさ、頻度を調べる。全体として男性形の翻訳の方が高得点を得る傾向があり、職業ごとの違いも固定観念的な性別表象に沿っていることが分かった。ただし、この選好の強さと一貫性は評価器と言語によって大きく異なる。本研究の結果は、MT評価には依然としてジェンダーバイアスが存在するが、その程度を捉えるには単一の集約指標を超え、評価器の挙動の相補的な側面を見る必要があることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Gender bias remains a persistent concern in machine translation (MT), affecting both generated translations and their automatic evaluation. When a source text leaves a person's gender unspecified, translations may realize that person using masculine or feminine forms, and both MT systems and evaluation metrics may exhibit systematic preferences between these alternatives despite the source providing no basis for such a distinction. We study this behavior in the WMT 2026 Automated Translation Quality Evaluation Systems Shared Task using an occupation-balanced subset of GAMBIT+. We consider seven English-source language pairs, six from the original dataset, targeting Arabic, Czech, Greek, Icelandic, Russian, and Ukrainian, and extend the original resource with German. The subset contains 1,308 masculine/feminine translation pairs per target language, with three examples for each of the 436 ISCO-08 occupational groups. We evaluate shared-task submissions and baselines for score prediction and error annotation, examining the direction, magnitude, and frequency of gender-related differences. We find an overall tendency for masculine translations to receive higher scores, as well as differences per occupation following stereotypical gender representations, although the strength and consistency of this preference vary considerably across evaluators and languages. Our results show that gender bias remains present in MT evaluation, but that capturing its extent requires looking beyond a single aggregate measure to complementary dimensions of evaluator behavior.
著者のコメント
Accepted for publication at the 11th Conference of Machine Translation (WMT26), co-located with EMNLP 2026
arXiv ID: 2609.21490 / 要約の誤りについて