arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

英語から東シリア語への統計翻訳モデルと対訳データを作る

Machine Translation between English and Syriac (East Syriac Dialect) using Statistical Machine Learning

Hadiana Sliwa and Hossein Hassani

この論文をやさしく読む

ひとことで言うと

対訳データの少ない東シリア語について、聖書から英語との対訳を整備し、句に基づく統計翻訳モデルを学習した研究です。

何に役立つ?

公開されるデータ・スクリプト・モデルは、この言語の翻訳研究の比較基準や再現実験に使えます。考えられる用途は、低資源言語向けの言語処理研究の基盤整備です。

この研究の面白いところ

PDFからの文抽出だけで済ませず、3人が対応関係を確認し、文字上のばらつきも処理しています。自動指標に加え、11人の母語話者による適切さ・流暢さの評価も行っています。

どこまで分かった?

学習データは聖書で、主に英語からアッシリア語へのモデルです。日常会話や他分野の文章で同じ性能が出るとは要旨では示していません。初という主張や従来研究の不足は著者の位置づけで、ここで網羅的に確認したものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

UNESCOはアッシリア語(シリア語)を消滅の危機にある言語と位置付けている。アッシリア人は世界各地でこの言語を話しているが、話者人口は50万~150万人と不確かである。シリア語は自然言語処理でも最も研究の少ない言語の1つである。過去10年間の機械翻訳の進歩にもかかわらず、公開コーパスの不足と、特にMadnkhaya文字に見られるシリア文字の正書法上の複雑さのため、この言語は計算言語学の文献で全く顧みられてこなかった。 本研究では、Mosesの枠組みを用いて、英語からアッシリア語への初の句に基づく統計的機械翻訳(SMT)モデルを開発する。英語とシリア語の聖書全体から38,847組の対訳文を作成した。既存の新約聖書データセットに、PDF抽出から新規に構築した旧約聖書を統合し、専用の分割スクリプトと3人のバイリンガル注釈者による手動の対応確認を用いた。学習前に、シリア語側のコーパスで発音区別符号を除き、Byte-Pair Encodingで分割して、正書法に由来するデータの疎らさを減らす。 構成、データ分割の比率、言語モデルの次数、語順移動の上限、Operation Sequence Modelの使用有無を変えた6モデルを学習・評価した。最良構成の単語単位BLEUスコアは23.54である。アッシリア語母語話者11人による人手評価では、適切さと流暢さの平均点は5点満点中それぞれ3.42と3.34だった。この結果は、形態が豊かなセム語族の言語について聖書コーパスで学習した、同程度の低資源SMTモデルと整合する。コーパス、スクリプト、学習済みモデルは公開されており、この危機言語の今後の機械翻訳と幅広い自然言語処理に向け、初めて系統的に整備された英語・シリア語データセットと再現可能な比較基準を研究者に提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

UNESCO considers the Assyrian (Syriac) language an endangered language. Although Assyrians speak the language worldwide, the speaking population is uncertain (ranging from 500,000 to 1,500,000). Syriac is also one of the least studied languages in Natural Language Processing (NLP). Despite advances in Machine Translation (MT) over the past decade, the lack of publicly available corpora and the orthographic complexity of the Syriac script, specifically the Madnkhaya script, have left this language entirely ignored in the computational linguistics literature. This study develops the first phrase-based Statistical MT (SMT) model for English-to-Assyrian MT using the Moses framework. We created a dataset of 38,847 sentence pairs from the complete English and Syriac Bible, merging a pre-existing New Testament dataset with an Old Testament built from scratch through PDF extraction, using custom segmentation scripts and manual alignment review by three bilingual annotators. The Syriac side of the corpus undergoes diacritic removal and Byte-Pair Encoding tokenization to reduce orthographic sparsity before training. We trained and evaluated six models using different configurations and splitting-scheme ratios, language model order, distortion limits, and the inclusion of an Operation Sequence Model. The best-performing configuration achieves a word-level BLEU score of 23.54. Human evaluation by 11 native Assyrian speakers resulted in mean adequacy and fluency scores of 3.42 and 3.34 out of 5, respectively. These results are consistent with comparable low-resource SMT models trained on Biblical corpora for morphologically rich Semitic languages. The corpora, scripts, and trained model are publicly available, providing the research community with the first systematically curated English-Syriac dataset and a reproducible baseline for future MT and broader NLP work on this endangered language.

著者のコメント

17 pages, 4 figures, 8 tables

arXiv ID: 2609.18529 / 要約の誤りについて