arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

ベンガル語の医療固有表現抽出で多言語モデルと専用モデルを比較

Language Specificity vs. Domain Diversity: Benchmarking Transformers for Bangla Medical NER

Rakib Abdullah, Md. Maruful Islam Maruf

この論文をやさしく読む

ひとことで言うと

ベンガル語の医療文書から薬や症状などの名前を抽出するモデルを、3,179件のテストデータで比較した研究です。

何に役立つ?

ベンガル語の医療情報抽出でモデルを選び、どの表現の種類に改善が必要か判断する資料になります。診療上の意思決定の安全性を検証した研究ではありません。

この研究の面白いところ

ベンガル語専用モデルより、多言語のXLM-RoBERTaが良好で、最頻出の症状分類が最も難しいという結果を示しています。

どこまで分かった?

結果はこのテストセットと比較したモデル、プロンプト設定に基づきます。症状のF1値は0.4367であり、医療現場にそのまま導入できる精度を示すものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

資源の少ない言語での医療分野の固有表現抽出(NER)は、言語表現の多様さと専門領域の注釈付きコーパス不足のため難しい。本研究は、微調整した3つのTransformerエンコーダー、BanglaBERT、多言語BERT(mBERT)、XLM-RoBERTaを、ゼロショットおよび少数例プロンプト設定のGPT-4o miniと比較する、ベンガル語医療NERの包括的な実験ベンチマークを示す。大規模言語モデルを50件だけの部分集合で評価した先行研究と異なり、3,179件のテストセット全体で大規模に評価し、統計的に頑健で再現可能な基準値を提供する。微調整したXLM-RoBERTaはF1値0.5959を達成し、従来報告された最高値0.5848を超え、新たな最高水準となった。一方、言語専用のBanglaBERTはF1値0.4937で、多言語モデルを一貫して下回った。これは、高度に専門的な臨床分野では、事前学習における領域の多様性が言語の専用性より重要になり得ることを示す。表現の種類別の分析では、薬と専門医の分類はF1値0.83超と高い信頼性で認識されたが、症状の分類は学習データで最も頻度が高いにもかかわらずF1値0.4367で最も難しかった。さらに、微調整したTransformerモデルは最良のプロンプト設定の3.76倍の性能であり、資源の少ない言語の構造化された臨床情報抽出では、プロンプトのみの処理はなお不十分であると結論づける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Medical Named Entity Recognition (NER) for low-resource languages remains a challenging task due to high linguistic variability and a scarcity of domain-specific annotated corpora. This work presents a comprehensive empirical benchmark evaluating three fine-tuned transformer encoders-BanglaBERT, multilingual BERT (mBERT), and XLM-RoBERTa-against GPT-4o mini under zero-shot and few-shot prompting configurations for Bangla medical NER. In contrast to prior studies that evaluated large language models on limited subsets of only 50 samples, we conduct a large-scale evaluation across the full test set of 3,179 samples, providing statistically robust and reproducible baselines. Our fine-tuned XLM-RoBERTa model achieves an F1- score of 0.5959, establishing a new state-of-the-art and surpassing the previously reported best result of 0.5848. Crucially, we demonstrate that the language-specific BanglaBERT model consistently underperforms its multilingual counterparts with an F1-score of 0.4937, indicating that pretraining domain diversity can outweigh language specificity in highly specialized clinical settings. Furthermore, we present a detailed per-entity-type analysis for this task, revealing that Medicine and Specialist categories are recognized with high reliability, achieving F1- scores above 0.83, while the Symptom category remains the most challenging with an F1-score of 0.4367 despite being the most frequent training class. Finally, fine-tuned transformer models outperform the optimal prompting configuration by a factor of 3.76, confirming that prompt-only pipelines remain inadequate for structured clinical entity extraction in low-resource language environments.

著者のコメント

6 pages, 2 figures. Accepted at the 2026 IEEE International Conference on Biomedical Engineering, Computer and Information Technology for Health (BECITHCON), Dhaka, Bangladesh

arXiv ID: 2609.29101 / 要約の誤りについて