arXiv論文メモ
新着一覧
cs.AI / cs.CL / cs.LG · 査読状況未確認

小児トリアージを答える公開LLMの属性による判定変化を監査

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang

この論文をやさしく読む

ひとことで言うと

症状を変えずに患者の属性や状況だけを変えたとき、LLMによる小児救急の優先度判定がどれだけ変わるかを比べています。

何に役立つ?

臨床導入を検討するモデルについて、属性に不適切に左右されるリスクを事前に比較する監査に役立ちます。実際の診療での利用を推奨する結果ではありません。

この研究の面白いところ

大きいモデルや医療向けモデルほど属性への感度が低いとは限りません。全体の変化率だけでは、危険側に変わるかどうかや共通する失敗が見えにくい点も調べています。

どこまで分かった?

臨床事例文を用いた反実仮想評価であり、患者の転帰を調べた臨床試験ではありません。属性を変えても答えが変わりにくいことだけで、元のトリアージ判断の正確さや安全性が保証されるわけではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

救急部門(ED)のトリアージは重要性の高い優先順位付けであり、人口統計的情報、社会経済的情報、医療システムの状況が、重症度の割り当てに不適切な影響を与える可能性がある。オープンソースの大規模言語モデル(LLM)は、ローカルで動作しプライバシーを保つ臨床意思決定支援の候補として検討されることが増えている。しかし、反実仮想的なバイアスが、モデル系列、規模、医療分野モデル、分野適応済みモデルの間でどう異なるかは不明である。 本研究では、小児のEmergency Severity Index(ESI)予測を対象に、10種類のオープンソースLLMの反実仮想監査を比較する。実際の症例およびハンドブック形式の臨床事例文から出発し、臨床像を固定したまま、付加する人口統計、社会経済、医療アクセス、行動、社会状況、システム状況の変数を1つだけ変えた、対になる反実仮想の事例を作る。対象にはQwen2.5-7B、Qwen2.5-14B-Instruct、QLoRAで微調整したQwen2.5-7B、MedGemmaの各モデル、MedLLaMA2-7B、GPT-OSS-20B、GPT-OSS-120Bが含まれる。何らかの判定変化、過小トリアージ、過大トリアージ、ESIで1段階を超える変化、平均変化、平均絶対変化を測定する。 反実仮想への感度は大きく異なり、モデル規模の拡大や医療分野での事前学習によって一貫して低下するわけではなかった。微調整したQwen2.5-7Bが全体として最も低い感度を示し、何らかの変化が生じる率は5.27%、平均絶対変化は0.0534であった。基礎モデルでは、それぞれ16.02%と0.1706であった。複数の大規模モデルや医療分野モデルでは、より大きな変化が見られた。層別解析と相関解析により、集計された率に隠れる、臨床上重要な変化の方向と共通の失敗パターンも明らかになった。これらの結果は、臨床導入前にオープンソースLLMの公平性リスクを比較する、軽量で臨床的に解釈しやすい枠組みとして反実仮想監査を支持する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.

arXiv ID: 2610.01963 / 要約の誤りについて