arXiv論文メモ
新着一覧
cs.CL / cs.LG · 査読状況未確認

翻訳向け追加学習で一般能力を守っても指示追従は残らない

Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following

Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney

この論文をやさしく読む

ひとことで言うと

翻訳モデルの追加学習で一般能力を保つ方法が、訳文の丁寧さや文法上の性を指定する能力も保つか調べた。

何に役立つ?

考えられる用途は、指示に従う翻訳モデルの追加学習方法の選択である。要旨ではLlamaの2規模と複数の言語対で比較した。

この研究の面白いところ

一般ベンチマークで忘却を抑えても、翻訳固有の指示制御は維持されるとは限らないことを示す。

どこまで分かった?

EWCは一般能力の低下を抑えたが翻訳制御は保てず、制御課題を混ぜた方法の利点も未見の指示文へは移らなかった。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

対訳データで大規模言語モデルを追加学習すると翻訳品質は改善するが、破局的忘却が起きることがある。忘却を抑える方法は通常、一般的なベンチマークで能力をどれだけ保つかによって評価される。本研究は、その知見が機械翻訳向けの追加学習や、敬体・常体、文法上の性、長さの制御など、訳文を変える翻訳固有の指示への追従(MT-IF)にも当てはまるかを問う。 補助データ、モデルの出力、基盤モデルのパラメータにそれぞれ基づく方法を比較した。まずLlama 3.2 1B Instructで選別調査を行い、次にアラビア語と英語、またはスペイン語と英語の双方向データで追加学習したLlama 3.1 8B Instructを評価した。Elastic Weight Consolidation(EWC)は両段階で一般能力を最もよく保った。8Bのスペイン語モデルでは、一般ベンチマークの平均スコア低下が通常の追加学習で11.0ポイントだったのに対し、EWCでは1.7ポイントだった。しかし、丁寧さと文法上の性の制御のスコアは通常の追加学習とほぼ同じ水準にとどまった。これらの制御を保ったのは制御課題の例を混ぜて学習する方法だけだったが、その利点は同じ課題の未見の指示文には移らなかった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Fine-tuning large language models on parallel data improves translation quality but can cause catastrophic forgetting. Mitigation methods are generally evaluated by retention on general benchmarks. We ask whether these findings transfer to machine translation (MT) fine-tuning and to MT-specific instruction following (MT-IF): instructions that modify a translation, such as formality, grammatical gender, and length control. We compare methods anchored to auxiliary data, to model outputs, and to the base model parameters, first in a screening study with Llama 3.2 1B Instruct, then on Llama 3.1 8B Instruct fine-tuned on bidirectional Arabic-English or Spanish-English data. Elastic Weight Consolidation preserves general capabilities best in both stages; on the 8B Spanish model the average score on general benchmarks drops 1.7 points versus 11.0 for standard fine-tuning, yet its scores for formality and grammatical gender control remain close to standard fine-tuning. Only data mixing with control-task examples preserves these controls, but its gains do not transfer to unseen prompts for the same task.

著者のコメント

Accepted at WMT 2026

arXiv ID: 2609.28395 / 要約の誤りについて