arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

救急トリアージでLLMが初期の訴えに引きずられる問題

LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage

Dipankar Srirag, Haokai Zhao, Ashutosh Kumar, Eleanor Hopper, Michael Dalton, Quoc Dung Nguyen, Aditya Joshi, Salil S. Kanhere, Padmanesan Narasimhan

この論文をやさしく読む

ひとことで言うと

救急患者との会話を途中まで読んで緊急度を決める課題で、LLMが後から出る情報を十分に反映できるか調べた研究。

何に役立つ?

救急トリアージ用LLMを評価する際、完成した記録での成績だけでなく、会話途中での判断を検証する必要性を示す。

この研究の面白いところ

モデルは後半の臨床情報を抽出できていても、判断は初期の主訴に引きずられ、専門臨床家との成績差が大きかった。

どこまで分かった?

評価したのは6モデル、生成会話425件と医師作成会話50件、5時点である。特定のデータセットでの結果であり、すべてのLLMの性能を断定するものではない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

救急部門のトリアージは、会話の進行に応じて判断する逐次的な過程である。既存の大規模言語モデル(LLM)評価は、完成した過去の記録を用い、医師に近い成績を報告している。本研究は、看護師と患者の会話を途中まで読んだ時点で緊急度ラベルを予測する、逐次トリアージの評価方法を実装した。LLMが生成した会話425件と、医師が作成した会話50件の2つのコーパスを用い、いずれもEmergency Severity Index(ESI)でラベル付けした。6つのLLMを、会話の途中の5時点で評価した。 二次重み付きκ係数(QWK)で測ると、すべてのモデルは完成した記録では中程度からかなり高い一致を示すのに対し、各途中時点では低めから中程度の一致へ低下した。制御した入力変更では、各時点でのラベルが最初の主訴のやり取りに引きずられ、プロンプトによる介入でもこの頭打ちは改善しなかった。モデルは後半の発話から臨床上重要な内容を抽出しているが、時点が進むにつれて真のラベルに対する意外性が増し、証拠を統合できていない。同じ会話を評価した専門臨床家3人のQWKは0.887~0.929だったが、最良のモデルは0.295だった。モデルの予測はESI-2とESI-3に集中し、正解との一致よりモデル同士の一致が高いため、アンサンブル化は失敗を悪化させた。完成済み記録による評価だけに基づいて救急トリアージにLLMを導入すると、この逐次的な失敗を見落とす。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-24 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.

著者のコメント

Under Review

arXiv ID: 2609.22904 / 要約の誤りについて