医療要旨を分類するモデルの説明手法の一致度を調べる
A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
この論文をやさしく読む
ひとことで言うと
医療文書の分類理由が、説明の作り方によってどれだけ変わるかを、5つの手法で比較しています。
何に役立つ?
分類モデルの説明を1つだけ信用せず、複数手法の一致度や誤りの傾向から監査するために役立ちます。
この研究の面白いところ
予測が曖昧なカテゴリでは、説明同士の一致も崩れることを調べています。語彙に過敏な反応など、説明の不一致と関わる失敗を整理しています。
どこまで分かった?
対象はDeBERTa-v3による医療要旨の分類で、患者の診断や治療の有効性を検証したものではありません。要旨には正解率や一致度の具体的な数値はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
医療要旨のゼロショット分類におけるDeBERTa-v3を監査するため、説明可能性を比較する枠組みを示す。同じ入力と予測に対して、異なる寄与度推定法が食い違う説明を生むという、説明可能な人工知能における不一致問題を扱う。Medical Abstractsコーパスに対して自然言語推論エンジンを実装し、診断カテゴリごとに内容を充実させた5つの仮説と、各クラス1,000文書の均衡した標本を用いる。 5つの説明法を比較する。モデルに依存しない方法としてSHAPとLIME、深層学習に特化した方法として遮蔽とInput × Gradient、Transformerに特化した方法としてAttention × Gradientである。寄与度の高いトークンによって説明を標準化し、手法の組ごとの一致度をJaccard指数で定量化する。 明確に定義された臨床領域では高い予測精度を達成する一方、意味的な曖昧さが大きい条件では性能が低下する。説明の安定性は予測の確かさを直接反映し、意味が一意に定まるカテゴリでは強い一致を示すが、診断上の不確実性があると顕著に低下する。さらに、誤りの定性的な監査から、語彙への過敏性、意味の重なり、寄与度説明の整合性の喪失という3つの系統的な失敗機序が見つかる。結果は、医療文書分類のTransformerモデルを監査する際、複数の説明法と定量的な一致指標を組み合わせることを支持し、広い診断ラベルよりも具体的な臨床オントロジーを優先することを示唆する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
A comparative explainability framework is presented to audit DeBERTa-v3 under zero-shot classification of medical abstracts. The work addresses the disagreement problem in Explainable Artificial Intelligence, where different attribution methods produce divergent explanations for the same input and prediction. A natural language inference engine is implemented over the Medical Abstracts corpus with five enriched hypotheses per diagnostic category and a balanced sample of one thousand texts per class. Five explanation methods are compared: SHAP and LIME as model-agnostic approaches, occlusion and Input x Gradient as deep-learning-specific approaches, and Attention x Gradient as a transformer-specific approach. Explanations are standardized through top-token attribution, and pairwise agreement is quantified using the Jaccard index. High predictive accuracy is achieved across well-defined clinical domains, whereas performance degrades under high semantic ambiguity. Explanatory stability directly mirrors predictive certainty, exhibiting strong convergence in univalent categories and a marked drop under diagnostic uncertainty. Furthermore, qualitative error auditing uncovers three systemic failure mechanisms: lexical hypersensitivity, semantic overlap, and loss of attribution coherence. The results support the combined use of several explanation methods and quantitative agreement metrics when auditing transformer-based models in medical text classification, and suggest prioritizing specific clinical ontologies over broad diagnostic labels.
著者のコメント
18 pages, 6 figures
arXiv ID: 2610.02116 / 要約の誤りについて