arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

人による修正訳を使いローカライズの翻訳品質評価を改善

LocQE: Principled Domain Adaptation for Localisation Quality Estimation by Leveraging Post-Edits

Kathy Hämmerl, Gabriel Bretschner, Joern Wuebker

この論文をやさしく読む

ひとことで言うと

一般的な翻訳評価モデルを、人が修正した訳のデータで調整し、数字・空白・句読点などが重要なローカライズでも使いやすくする研究です。

何に役立つ?

同じ原文に対する複数の候補訳から、実際に採用される修正訳を選ぶ品質評価の改善に役立ちます。

この研究の面白いところ

修正訳の方が好まれるという順位情報だけでなく、連続的なスコアも組み合わせ、絶対評価と候補間比較を一緒に安定化させています。

どこまで分かった?

要旨には学習データの具体的な件数や改善幅が示されていません。未見領域への弱さを扱う研究ですが、あらゆる言語・ローカライズ対象で有効だという保証は記されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

COMETKiwiのような学習済み品質推定(QE)モデルは広く使われ、一般的な機械翻訳評価ではよく機能する。しかし、未見の領域では苦戦することが知られており、実際のローカライズでの性能が制約される。本研究では、数字が正確に訳されているか、さらには空白や句読点の個数が正しく保たれているかなど、ローカライズで重要な要因の一部に対して、これらのモデルが鈍感であることを示す。また、機械翻訳の最適化に必要な重要能力として、同じ区間の異なる翻訳を正しく順位付けする能力があるが、これも領域の移行により大きく損なわれる。 大規模な直接評価データがない状況で、少量のポストエディットデータでも領域間の隔たりを減らすため、原則に基づくファインチューニング手法を提案する。マルチタスクのファインチューニングと単純なトークナイザーへの介入を用い、ローカライズにおいて、好まれる修正後の訳と、採用されなかった初期訳を、明らかによりよく区別できるQEモデルを作成する。選好と人工的な連続スコアが互いを安定化することを示し、絶対スコアと同じ原文の翻訳間の比較という両面で評価指標を較正するには、両種類の信号が必要だと論じる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Learned quality estimation (QE) models such as COMETKiwi are widespread and work well for general machine translation evaluation. However, they are known to struggle on unseen domains, limiting their performance in a real-world localisation context. We show that they are insensitive to some important factors in localisation, such as whether numbers are translated accurately, or even whether the correct number of spaces and punctuation are preserved in a translation. Further, a key capability for optimisation of machine translation is the ability of QE models to accurately rank different translations of a single segment, which suffers significantly from the domain transfer. In the absence of large-scale direct assessment data, we propose principled fine-tuning approaches to reduce the domain gap with even small amounts of post-editing data. Using a multi-task fine-tuning approach and a simple tokeniser intervention, we create a QE model which proves markedly better at distinguishing preferred post-edits from rejected initial translations in a localisation context. We show that preferences and artificial continuous scores stabilise each other, and argue that to calibrate metrics both in terms of their absolute scores and comparisons between translation of the same source, both types of signal are needed.

arXiv ID: 2609.18720 / 要約の誤りについて