がん治療支援LLMの過剰な拒否を合成症例で評価する
OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation
この論文をやさしく読む
ひとことで言うと
治療案を厳しく拒否するだけで安全と評価してよいかを、合成した肺がん症例で調べています。根拠が一部ある案と、根拠のない案を区別することが中心です。
何に役立つ?
治療支援AIを評価する際に、総合安全性スコアだけでなく、医師の確認を要する選択肢が適切に残っているかを調べるための材料になります。
この研究の面白いところ
根拠確認や情報不足の判定を分離した方法では、過剰な拒否が6.7%となりました。医師間の不一致も調べ、ラベルの境界自体にある難しさを扱っています。
どこまで分かった?
500件の合成症例での評価であり、患者の治療成績を改善した臨床試験ではありません。正解率91.2%はこのベンチマークの分類結果です。医師によるラベル付けの検討は2人で行われています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
分子腫瘍ボードは、精密がん医療を支えるため、ゲノム所見、臨床的な状況、治療の根拠を統合する。AIをこの作業に取り入れる際の重要な安全上の課題は、本当に根拠のない推奨と、情報不足、ECOGによる全身状態の不良、その他の臨床上の注意点から腫瘍医の確認が必要ではあるものの、根拠に支えられた選択肢とを区別することである。 本研究は、5種類の敵対的な誤りの分類と、「支持される」「部分的に支持される」「支持されない」「情報不十分」の4つの安全性ラベルを備えた、500件の合成非小細胞肺がん症例からなるオープンソースのベンチマークOpenMTB-Auditを導入する。大規模言語モデルの8構成を調べたところ、過剰な拒否が広く見られた。すべての構成で、真に「部分的に支持される」症例の83.3〜100%について、そのラベルを維持できなかった。臨床に即して調整された推論ではなく、ラベルの区別が失われることで、集約された安全性スコアが高くなっていた。 この問題に対処するため、根拠の検証、欠けている情報の検出、安全性の分類、判断の保留を分離した、決定論的な7モジュールの枠組みMTB-AuditAgentを開発した。この方法は過剰な拒否を6.7%に減らし、91.2%の正解率を達成した。95%信頼区間は88.6〜93.6%である。腫瘍医2人によるアノテーション研究では、情報の十分さと治療の最適化の境界に意見の不一致が集中しており、臨床上意味のある区別を保つ必要性が示された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from evidence-supported options that still require oncologist review because of incomplete information, poor ECOG performance status, or other clinical caveats. We introduce OpenMTB-Audit, an open-source benchmark of 500 synthetic non-small cell lung cancer cases spanning five adversarial error categories and four safety labels: Supported, Partially Supported, Unsupported, and Insufficient Information. Across eight large language model configurations, we identify pervasive over-refusal: all LLM configurations failed to retain the Partially Supported label in 83.3-100% of true Partially Supported cases, achieving high aggregate safety scores through label collapse rather than clinically calibrated reasoning. To address this limitation, we developed MTB-AuditAgent, a deterministic seven-module framework separating evidence verification, missing-information detection, safety classification, and abstention. It reduces over-refusal to 6.7% and achieves 91.2% accuracy (95% CI: 88.6-93.6%). A two-oncologist annotation study found disagreement concentrated at the boundary between information sufficiency and treatment optimization, underscoring the need to preserve clinically meaningful distinctions.
著者のコメント
Accepted for oral presentation and publication at the Pacific Symposium on Biocomputing (PSB) 2027
arXiv ID: 2610.01497 / 要約の誤りについて