ヨルバ語の声調などの記号を小型モデルで復元する
Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá
この論文をやさしく読む
ひとことで言うと
ヨルバ語の文章で省略された声調などの記号を、バイト単位の言語モデルで補う研究です。
何に役立つ?
記号が省かれた文章を自然言語処理に使う前の整備や、意味の曖昧さを減らす補助への利用が考えられます。
この研究の面白いところ
比較対象のmT5-baseと同等の復元性能を、およそ半分のパラメータ数で得たと報告しています。復元以外の文章内容を忠実に保つ点も評価しています。
どこまで分かった?
具体的な数値はYADベンチマークで統一手順により測った結果です。著者ら自身が、より大規模で用途に合ったベンチマークの整備を求めています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ヨルバ語は広く話されている声調言語であり、語彙の曖昧さを避けるためにダイアクリティカルマーク(発音区別符号)に依存する。しかし、これらの記号を省略して書かれることが多く、後続の自然言語処理(NLP)の課題を妨げている。 本論文では、ByT5-smallを微調整した、バイト単位の自動発音区別符号復元(ADR)モデルYo-ByT5を導入する。統一した手順の下で、公開済みのヨルバ語ADRモデル五つと、重みが公開されている大規模言語モデル(LLM)一つとともに、YADベンチマークでYo-ByT5を評価する。 結果から、Yo-ByT5はDERが10.14%、CERが3.48%となり、既存の最も強いモデルであるmT5-baseと同等の性能を示した。さらに、パラメータ数がmT5-baseのおよそ半分であるにもかかわらず、より高い文章の忠実性を示した。学習コードとモデル出力も公開し、ヨルバ語の発音区別符号復元に特化した、より大規模なベンチマークの開発を呼びかける。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.
著者のコメント
7 pages, 3 figures, 3 tables. Code and outputs: https://github.com/lazy-monster/yo-byt5
arXiv ID: 2610.01634 / 要約の誤りについて