arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

手術動画で文章が指す器具を切り出す大規模評価データ

LD-RSVIS: A Large-Scale and Diverse Benchmark for Referring Surgical Video Instrument Segmentation

Zan Wang, Yunhe Feng, Dong Nie, Oluwatosin Oluwadare, Kewei Sha, Yan Huang, Heng Fan

この論文をやさしく読む

ひとことで言うと

手術動画で文章が指す器具を切り出す手法を評価するため、対象がない場合や複数ある場合も含む大規模データを作った研究。

何に役立つ?

手術動画の器具切り出し手法を、多様な器具や指示文で学習・比較するための評価基盤になる。要旨では12の既存手法と提案手法を評価している。

この研究の面白いところ

3536動画・109万フレームを手作業で注釈し、従来扱われにくかった対象なしと複数対象の指示も含めた。

どこまで分かった?

提案データとコードは要旨では公開予定とされている。提案手法の性能は有望と述べるが、数値指標や実際の手術支援での有効性は要旨に示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文章で指定された器具を手術動画から切り出す課題はRSVISと呼ばれる。研究が進む一方、従来のモデルは比較的小さな評価データで学習・評価されることが多く、より一般的な手法の開発を妨げている。既存のデータは動画内の一つの器具を指す表現だけを扱い、複数の器具を指す表現や該当器具がない表現を見落としているため、実際の場面での利用も制限される。本研究では、より頑健で一般的なRSVISを促す評価データLD-RSVISを提案する。25種類の手術から得た3536本の動画、計109万フレームを含み、30種類の器具を対象とする。動画と器具の種類が多いため、一般的な手法の大規模な学習と評価に役立つ可能性がある。 既存データと異なり、対象なし、一対象、複数対象という多様な参照設定を含む。注釈の品質を保つため、すべての動画を手作業でラベル付けし、複数回の点検と修正を行った。著者らの知る限り、RSVISでは最大規模かつ最も多様な評価データである。比較の基準を示すため代表的な12手法を評価したところ、なお改善が必要と分かった。また、文章に含まれる互いに補完的な複数の手掛かりから対象固有の情報を取り出し、その情報と文章を使って切り出す、簡潔なCascade-RSVISも示し、有望な性能を得た。評価データとコードは公開予定である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Referring surgical video instrument segmentation (RSVIS) aims at segmenting the instrument in a surgical video, given a textual description. Despite recent progress, current models are trained and assessed on relatively small-scale benchmarks, hindering the development of more general RSVIS. In addition, existing benchmarks support only the single-target expression that refers to one instrument in the video, while overlooking multi-target and no-target referring expressions, restricting the applicability of RSVIS in practical scenarios. Addressing these issues, we propose LD-RSVIS, a new benchmark aiming to facilitate more robust and general RSVIS. Specifically, LD-RSVIS consists of 3,536 surgical videos with 1.09 million frames and covers a broad set of 30 instrument classes from 25 various procedures. By including abundant videos and classes, LD-RSVIS could benefit large-scale training and evaluation of more general RSVIS methods. Besides, unlike existing datasets, LD-RSVIS offers diverse referring settings, including no-target, single-target, and multi-target expressions, which enables the development of more practical RSVIS models in real applications. In order to ensure high-quality annotations, all videos in LD-RSVIS are manually labeled with multiple rounds of inspection and refinement. To our knowledge, LD-RSVIS is the largest and most diverse benchmark for RSVIS. To analyze LD-RSVIS and to provide comparison for future research, we evaluate 12 representative methods, and the results reveal that more efforts are required for improvements. To encourage future research, we present a simple yet effective RSVIS method, dubbed Cascade-RSVIS, that first mines target-specific cues using the complementary multi-cue text information and then employs such cues and textual information for segmentation, achieving promising performance. Our benchmark and code will be released.

arXiv ID: 2609.23067 / 要約の誤りについて