レバント方言アラビア語の発音を音声で評価するデータセット
SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic
この論文をやさしく読む
ひとことで言うと
五つのレバント方言変種について、音声と三種類の表記を対応させ、発音関連の処理を比べられる1300発話のデータセットです。
何に役立つ?
音声認識や発音への変換が、どの方言変種でうまくいくかを分けて評価するのに役立ちます。
この研究の面白いところ
標準化が不十分な文字表記だけに正解を依存させず、音声を基準に複数の注釈層を対応付けています。
どこまで分かった?
対象は選ばれた五つの変種と1300発話です。要旨にはモデル別の成績や、ほかの地域変種への評価結果は載っていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
レバント方言アラビア語(LA)は数千万人が話しており、その音声・言語技術を評価する共通ベンチマークが強く求められている。LAには内部的な多様性があり、正書法から発音が分かりにくく、表記も標準化されていないため、その技術の評価は特に難しい。 本研究は、公開音声コーパスから抽出した1300発話からなるベンチマークSHAMS(SHami Annotated Multi-dialect Speech)を提示する。パレスチナの都市・農村変種と、ヨルダン、レバノン、シリアの都市変種という五つのLA変種を均等に含む。各発話は、音声、母音記号のない表記、発音区別符号を付けたテキスト、音声転写という、対応付けた四つの層で表現する。 この構造により、発音区別符号付与、書記素から音素への変換、自動音声認識、音声から音素への変換などの下流課題を、音声に根拠を置き、変種別に分けて評価できる。これらの課題で公開モデルとプロプライエタリモデルを評価し、LA全体の進歩を測るうえでの本ベンチマークの有用性を示す。SHAMSは https://shams-nlp.github.io で公開する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotated Multi-dialect Speech), a benchmark comprising 1,300 utterances drawn from open audio corpora, balanced across five LA varieties (Urban and Rural Palestinian, and Urban Jordanian, Lebanese, and Syrian). Each utterance is represented across four aligned tiers: audio, unvocalized orthography, diacritized text, and phonetic transcription. This structure supports evaluation of various downstream tasks such as diacritization, grapheme-to-phoneme conversion, automatic speech recognition, and audio-to-phoneme, grounded in audio and stratified by variety. We benchmark open and proprietary models across these tasks to demonstrate the utility of this benchmark for measuring progress across LA. We release SHAMS at https://shams-nlp.github.io .
著者のコメント
Accepted to ArabicNLP 2026. Project page: https://shams-nlp.github.io/
arXiv ID: 2610.01427 / 要約の誤りについて