arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

胆嚢摘出術の安全視野を評価する多段階画像モデル

CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment

Bock-Zien Toh, Yuanchuan Ren, Tay Aw Yu, Ng Khee Ong, Zhehua Mao, Sophia Bano

この論文をやさしく読む

ひとことで言うと

胆嚢摘出術の画像で解剖構造を検出・分割し、安全視野の基準を評価する。

何に役立つ?

考えられる用途は、手術映像の安全視野の確認を支援する研究である。臨床判断を代替する検証ではない。

この研究の面白いところ

検出した枠とマスクを次の課題に使い、段階的な学習で不安定さを抑える。

どこまで分かった?

Endoscapesの未使用テスト集合で改善を示した。臨床現場での安全性や転帰は要旨に記載がない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

腹腔鏡下胆嚢摘出術におけるCritical View of Safety(CVS)の自動評価には、三つのCVS基準の認識と、小さくまれで隠れやすい肝胆嚢周辺の構造を正しく捉えることの両方が必要である。従来の学習法は画像全体の分類、検出、領域分割、グラフ推論など、使う解剖情報が異なるが、安全上重要な構造の位置付けが主要な障害である。提案するCasCVS-Netは、物体検出、意味領域分割、CVS評価を一緒に行う段階的な多課題モデルで、Endoscapesデータセットで学習する。予測した枠が領域分割を導き、予測したマスクがCVS分類の領域特徴を与える。推論時のCVS評価では、正解注釈ではなくモデルの予測だけを使う。結合した学習の不安定さを抑えるため、検出、検出と領域分割、三課題全体へと段階的に学習し、最後に課題別に微調整する。公開の未使用テスト集合では、三課題すべてで対応する単一課題の比較手法より良く、検出mAP 32.0、意味領域分割mIoU 46.8、まれな解剖構造のmIoU 15.3、CVS mAP 67.2を得た。最新手法のLG-CVSとSV2LSTGに対して、CVS mAPは相対値でそれぞれ6.3%、4.5%高く、絶対値では4.0、2.9ポイントの差だった。予測した枠とマスクを介して課題を結び付ける方法が、とくにまれな肝胆嚢構造の位置付けを改善すると示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Automated assessment of the Critical View of Safety (CVS) in laparoscopic cholecystectomy requires both recognition of the three CVS criteria and anatomical grounding in small, rare, and often occluded hepatocystic structures. Learning-based methods differ in the anatomical information they use, from image-level classification to detection, segmentation, or graph-based reasoning, yet grounding the safety-critical anatomy remains the main bottleneck. We propose CasCVS-Net, a staged multi-task cascade that jointly performs object detection, semantic segmentation, and CVS assessment, trained on the Endoscapes dataset. The model couples the tasks through predicted anatomy: predicted boxes guide segmentation, and predicted masks provide region-level features for CVS classification, so CVS assessment at inference uses only model predictions rather than ground-truth annotations. To reduce optimisation instability in this coupled setting, training progresses from detection to detection-segmentation and then to the full three-task cascade, followed by task-wise fine-tuning. Evaluation on the public unseen test set shows that CasCVS-Net improves over matched single-task baselines on all three tasks, achieving 32.0 detection mAP, 46.8 semantic mIoU, 15.3 rare-anatomy mIoU, and 67.2 CVS mAP. It outperforms the state-of-the-art LG-CVS and SV2LSTG by 6.3% and 4.5% relative CVS mAP, respectively, corresponding to 4.0 and 2.9 mAP points. These results show that staged task coupling through predicted boxes and masks improves anatomical grounding for CVS assessment, particularly for rare hepatocystic structures.

著者のコメント

10 pages, 3 figures, 3 tables

arXiv ID: 2609.27681 / 要約の誤りについて