arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

機械翻訳を介した会話の「伝わりやすさ」を評価

Evaluating Communicative Success in Machine-Translated Conversation

Faiz Ghifari Haznitrama and Alice Oh

この論文をやさしく読む

ひとことで言うと

翻訳された会話が、意味の正しさだけでなく意図や文化的な適切さまで伝えているか評価する枠組みです。

何に役立つ?

通訳エージェントの会話品質を、文単位の翻訳指標が見落とす失敗も含めて検査するために使えます。単発発話と複数ターンの会話を扱います。

この研究の面白いところ

アラビア語、ベンガル語、インドネシア語、韓国語の5,624場面で10構成を比較し、意味、語用、文化・社会面の順に成功が低下しました。評価法自体を人の注釈や制御された変更でも点検しています。

どこまで分かった?

複数ターンでは模擬利用者が応答する設定を含みます。文脈や構造化指示の効果は構成ごとに異なり、要旨にはすべての実会話での効果を保証する結果はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

機械翻訳(MT)に基づく通訳エージェントは、共通言語を持たない人々のリアルタイム会話を仲介する機会を増やしている。しかし、その評価には依然として孤立した文を対象に作られた指標が使われており、意思疎通の成否よりも原文への忠実度を測っている。本研究では、通訳を介した会話を意味、語用、文化・社会という3層で評価する、再利用可能なチェックリストと判定器の枠組みを導入する。忠実度の指標では測れない自然さ、意図、社会的な適切さを扱う。 この枠組みは、単一ターンと対話的な複数ターンの両設定で動作する。複数ターンでは、会話の進行に合わせて模擬ユーザーが翻訳されたメッセージに応答し、各ターンと会話全体を評価する。制御された摂動、判定器間の比較、人手アノテーションを通じて広範に検証した。主たる単一ターンのベンチマークでは、OpenSubtitles由来の5,624シナリオを使い、アラビア語、ベンガル語、インドネシア語、韓国語にまたがる12の翻訳方向で10通りの通訳構成を評価する。複数ターンの研究では、台本ありとライブの両モードで、全6言語ペアを対象とする。 結果として、意味上の成功から語用上、文化・社会上の成功へと一貫して成績が下がる。従来のMT指標は、性能の高い通訳で生じる失敗を見落とす。プロンプトの要素を除く比較実験では、シナリオの文脈、構造化された指示、文化的な文脈が意思疎通の成功を改善したが、改善幅は構成によって異なる。本研究は、会話における通訳エージェントの評価枠組みとベンチマークを提供し、既存の翻訳指標に加えて意思疎通の成功を評価する重要性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Interpreter agents built on machine translation (MT) increasingly mediate live conversation between people who do not share a language, yet we still evaluate them with metrics built for isolated sentences, which measure fidelity rather than whether communication succeeds. We introduce a reusable three-layer checklist-and-judge framework that evaluates interpreter-mediated conversation across semantic, pragmatic, and cultural-social dimensions, covering the naturalness, intent, and social appropriateness that fidelity metrics leave unmeasured. It runs in both single-turn and interactive multi-turn settings, where simulated users reply to translated messages as the conversation unfolds and each turn is scored alongside the conversation as a whole. We extensively validate it through controlled perturbations, cross-judge comparisons, and human annotations. Our main single-turn benchmark evaluates 10 interpreter setups across Arabic, Bengali, Indonesian, and Korean from 5,624 OpenSubtitles-derived scenarios spanning 12 translation directions, and our multi-turn study covers all 6 language pairs in scripted and live modes. Results show a consistent decline from semantic to pragmatic and cultural-social success, while conventional MT metrics overlook failures among stronger interpreters, and prompt ablations show that scenario context, structured instructions, and cultural context improve communicative success, although gains vary across setups. Our work thus provides an evaluation framework and benchmark for interpreter agents in conversation, and highlights the importance of communicative success alongside existing translation metrics.

著者のコメント

32 Pages, 11 Figures, 11 Tables

arXiv ID: 2609.19885 / 要約の誤りについて