がん登録担当者を支援する文献検索付き対話システム
CRISS: A Retrieval-Augmented AI Chatbot for Assisting Cancer Registrars
この論文をやさしく読む
ひとことで言うと
がん登録の基準を検索して根拠を引用する対話システムを作り、RAGなしの回答より根拠との整合性が高いと報告した。
何に役立つ?
がん登録担当者が更新される符号化・病期分類の指針を探す際の支援が考えられる。
この研究の面白いところ
質問の難易度別にRAGあり・なしを比べ、商用モデルとローカルモデルの違いも調べている。
どこまで分かった?
評価にはLLM判定者を用いた。最終的な登録情報の判断は人が担う設計で、実務上の誤り減少を直接示す数値は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
腫瘍データ専門家を含むがん登録担当者は、複雑で頻繁に更新される符号化基準と病期分類基準を解釈しなければならない。著者らは、登録業務の指針に根拠の引用付きで素早くアクセスできる、検索拡張生成(RAG)による対話型支援システムCRISSを開発した。本研究では、正確で引用に支えられた回答、関連指針へのアクセスと解釈の改善、そして最終的な情報抽出の判断を人が担う形での研修やヘルプデスク業務への支援が可能かを評価した。 全国的ながん登録基準から分野特化の知識ベースを作り、メタデータ付きの文章片に分割して密ベクトル表現として索引化した。検索された文章片を大規模言語モデルによる引用に基づいた回答の生成に利用した。Gemini系列とGPT系列の公開重みモデル、商用モデル、およびRAGなしの基準モデルを、易しい・中程度・難しい登録業務の質問で評価し、LLMを判定者に用いた。RAG構成は一貫してRAGなしの方法を上回り、特に難しい質問で差が目立った。根拠との整合性の平均スコアは、易しい・中程度・難しい質問の順にRAGで0.62、0.56、0.59、RAGなしでは0.29、0.26、0.29だった。RAGモデルは意味的類似度でも全体に高かった。商用RAGモデルは易しい・中程度の質問で最も良く、ローカルRAGモデルは難しい質問で最高順位となり、商用モデルは全般により慎重だった。 分野特化RAGは、がん登録の質問への回答の根拠付けと品質を改善し、難易度にまたがって引用付きの支援を可能にした。CRISSは、最終的な符号化判断を人が担いながら登録担当者を支援する、根拠に基づくAIの可能性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Cancer registrars, including Oncology Data Specialists (ODSs), must interpret complex and frequently updated coding and staging standards. We developed CRISS (Cancer Registry Intelligent Support System), a retrieval-augmented generation (RAG) conversational assistant that provides rapid, citation-supported access to registry guidance. This study evaluated whether CRISS could (1) support accurate and citation-supported responses, (2) improve access to and interpretation of relevant guidance, and (3) support training/helpdesk use while preserving human oversight of final abstraction decisions. We built a domain-specific knowledge base from national cancer registry standards, segmented into metadata-tagged passages and indexed as dense embeddings. Retrieved passages were used to generate citation-grounded responses through a large language model (LLM). Open-weight, proprietary, and non-RAG baseline models across Gemini and GPT families were evaluated on easy, medium, and hard registry questions using an LLM-as-a-Judge protocols. RAG configurations consistently outperformed non-RAG approaches, especially as question difficulty increased. Mean grounding scores for RAG were 0.62/0.56/0.59 across easy/medium/hard tiers versus 0.29/0.26/0.29 for non-RAG. RAG models also achieved higher semantic-similarity scores overall. Proprietary RAG models performed strongest on easy and medium questions, while local RAG models ranked highest on hard questions and proprietary models were generally more cautious. Domain-specific RAG improved evidence grounding and response quality for cancer registry questions while enabling citation-supported assistance across complexity levels. CRISS demonstrates the potential of human-centered, citation-grounded AI to support cancer registrars while preserving human oversight for final coding decisions.
著者のコメント
21 pages, 13 figures, 7 tables. Keywords: cancer registry, retrieval-augmented generation, large language models, conversational AI, clinical informatics, oncology data specialists, medical question answering, AI safety, clinical decision support
arXiv ID: 2609.29075 / 要約の誤りについて