arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

文脈依存機械翻訳の注意ヘッドを勾配で大規模分析

Scaling Attention Head Analysis via Gradient-Based Attribution in Context-Aware Machine Translation

Paweł Mąka and Yusuf Can Semerci and Jan Scholtes and Gerasimos Spanakis

この論文をやさしく読む

ひとことで言うと

翻訳モデルの注意ヘッドが曖昧さの解消にどう寄与するかを、損失の勾配から調べる手法である。

何に役立つ?

文脈依存の機械翻訳モデルで、どの注意ヘッドが性能に関係するかを大規模に調べるのに役立つ。

この研究の面白いところ

注意の平均量だけでは性能への寄与を説明できず、複数の関係で働く汎用的なヘッドも見つかった。

どこまで分かった?

評価は四モデル・四言語方向・50現象で行い、注意スコア操作との整合性はそのうち三モデル・二言語方向で確認した。すべてのモデルでの因果関係を検証したとは述べていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

トークン単位の最大マージン損失から注意マップへ勾配を逆伝播させる、注意ヘッドの寄与度評価法を導入する。この枠組みは注意ヘッドの大規模な因果的分析を可能にし、大規模言語モデルにも適用しやすい。文脈を考慮する機械翻訳の曖昧性解消を課題として、四つのモデルと四つの言語方向にわたる50種類の現象を分析する。三つのモデルと二つの言語方向では、特定のトークン間関係の注意スコアを増やしたときの効果と提案法の結果が整合することを実験的に示し、手法の頑健性を確認する。分析では、異なる関係に注意を向けた場合にもモデルの性能を改善する「汎用的な」注意ヘッドが見つかった。また、ある関係にヘッドが割り当てる平均的な注意の大きさは、必ずしもモデル性能と結び付かない。このことは、学習過程でヘッドの機能に冗長性が生じたことを示唆する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In this paper, we introduce a gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps. This framework enables a large-scale causal analysis of attention heads, making it suitable for LLMs. We evaluate our method on the task of disambiguation in Context-aware Machine Translation, where we analyze 50 phenomena across 4 models and 4 language directions. We empirically show the alignment of our method with the effects of increasing the attention scores of token-to-token relations on three models and two language directions, ensuring the robustness of our method. Our analysis reveals the presence of the "general-purpose" attention heads that improve the model's performance when attending to different relations. We find that the average attention a head assigns to a relation does not necessarily relate to the model's performance, which suggests that the models developed redundancies during training in terms of the head functions.

arXiv ID: 2609.28117 / 要約の誤りについて