arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

女性の健康相談を多言語と表現の違いから評価

HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication

Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi, Hana Essam Sayed Ahmed Amrya, Mariam Mousa

この論文をやさしく読む

ひとことで言うと

女性の健康に関する同じ相談を、言語や言い方を変えてAIが正しく理解できるかを評価します。

何に役立つ?

医療対話モデルの評価で、全体の正答率に隠れる過小評価や確認不足を見つける枠組みになります。

この研究の面白いところ

英語・フランス語・現代標準アラビア語で六つの表現形式を用意し、学習ラベルの言語間の不整合が危険な過小評価に関わると示します。

どこまで分かった?

再適応で過小トリアージは減っても、フランス語0.572、アラビア語0.558が残ります。これはモデル評価の結果で、臨床運用の安全性を保証するものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルは医療コミュニケーションで利用される機会が増えているが、多くの評価は、利用者の懸念を正しく解釈したと仮定したうえで、応答の質を重視する。本研究では、女性の健康に関するコミュニケーションの多言語理解を、条件を揃えて評価する枠組みHerHealthEvalを導入する。 各臨床事例について、英語、フランス語、現代標準アラビア語の対応する版を用意し、標準形、臨床的表現、一般の人の表現、間接的または断定を避ける表現、感情的な不安を示す表現、意図的に情報を不足させた表現という6つの形式を使う。最初の5つは同じ根本的懸念を表し、同じ臨床情報を保持する。一方、情報不足の形式は関連する詳細を意図的に省き、追加確認が必要だとモデルが認識するかを調べる。 多言語の指示対応モデルとQLoRAで適応した派生モデルを、懸念の分類、リスクの較正、追加確認の振る舞い、解析可能な出力形式への準拠、表現形式をまたぐ一貫性で評価する。結果は、集計した正解率や一貫性が、安全に関わる失敗を隠し得ることを示す。言語間で非対称なリスク教師情報を用いた多言語適応モデルでは、フランス語とアラビア語で過小トリアージの指標が0.994に達する。出典から導き、言語間で不変なリスクラベルを使って条件を制御した再適応を行うと、それぞれ0.572と0.558へ低下する。 これらの結果は、頑健な多言語医療評価には、使用域による表現の違い、不確実性への対応、適応用ラベルの由来と不変性を明示的に検査する必要があることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. We evaluate a multilingual instruction model and QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency. Results reveal that aggregate accuracy and consistency can conceal safety-relevant failures. A multilingual adaptation model reaches 0.994 under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using source-derived, language-invariant risk labels reduces under-triage to 0.572 and 0.558, respectively. These findings show that robust multilingual healthcare evaluation requires explicit testing of register variation, uncertainty handling, and the provenance and invariance of adaptation labels.

著者のコメント

8 pages, 2 figures, 3 tables. Submitted to the 2026 International Conference on Large Language Models (LLM 2026)

arXiv ID: 2609.20684 / 要約の誤りについて