手術の文脈と質問から器具の領域を切り出すSIRA
SIRA: Reasoning-Aware Surgical Instrument Segmentation via Query-Anchored Alignment
この論文をやさしく読む
ひとことで言うと
器具の種類だけでなく、手術の状況と質問の意味から画像内の対象領域を判断する研究です。専用データセットとモデルを提案します。
何に役立つ?
手術映像の文脈を踏まえた質問と、器具の画素単位の領域を結び付ける研究に役立ちます。ロボット支援や手術過程の分析が背景にあります。
この研究の面白いところ
対象の意味と質問が求める意味を分け、それぞれを視覚情報へつなぎます。41,000組の画像・テキストを用意しています。
どこまで分かった?
改善はSurgRS上の実験結果です。改善幅の数値、臨床での支援効果や安全性の検証は要旨には示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
手術器具セグメンテーション(SIS)はロボット支援や手術ワークフローの分析で重要な役割を果たす。しかし、既存のSIS手法の多くはセグメンテーションをカテゴリに基づく位置特定問題として定式化しており、手術ワークフローの手順の文脈やタスクに依存する意味を捉える能力が制限されている。本研究は、セグメンテーションを手術の文脈のもとで質問を条件とする推論として捉える、推論を考慮した手術器具セグメンテーション(RA-SIS)というタスクを導入する。この設定を評価するため、41,000組の画像・テキスト対からなる手術推論セグメンテーションデータセットSurgRSを構築する。個々の対象のマスクを構造化された質問・回答の教師情報と対応付け、意味を画素レベルに結び付けられるようにする。SurgRSに基づき、手術器具の推論・セグメンテーション支援システムSIRAを提案する。このマルチモーダル枠組みは対象レベルと質問レベルの意味を分離し、質問を基準とする二重の整合を通じて視覚特徴と統合する。質問の意味を空間特徴およびセグメンテーション用プロンプトと整合させることで、マスク予測における意味と視覚情報の一貫性を高める。SurgRSでの広範な実験は、既存の推論対応ベースラインに対する改善を示した。コードはhttps://github.com/linxir226/SIRAで公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Surgical instrument segmentation (SIS) plays a critical role in robotic assistance and surgical workflow analysis. However, most existing SIS methods formulate segmentation as a category-driven localization problem, limiting their ability to capture procedural context and task-dependent semantics in surgical workflows. We introduce Reasoning-Aware Surgical Instrument Segmentation (RA-SIS), a task formulation that frames segmentation as query-conditioned inference under surgical context. To benchmark this setting, we construct SurgRS, a surgical reasoning segmentation dataset consisting of 41,000 image-text pairs, which aligns instance-level masks with structured query-answer supervision to enable semantic grounding at the pixel level. Based on SurgRS, we propose Surgical Instrument Reasoning and Segmentation Assistant (SIRA), a multimodal framework that disentangles target-level and query-level semantics and integrates them with visual features through query-anchored dual alignment. By aligning query semantics with spatial features and segmentation prompts, SIRA enhances semantic-visual consistency in mask prediction. Extensive experiments on SurgRS demonstrate improvements over existing reasoning-aware baselines. Code is available at https://github.com/linxir226/SIRA.
arXiv ID: 2609.21402 / 要約の誤りについて