arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

動きと不確かさを考慮し、手術動画から技能を分類

Video-based Surgical Skill Assessment Using Dynamics-and-Uncertainty-Aware Tree-based Gaussian Process Classifier

Arefeh Rezaei, Mohammad Javad Ahmadi, Amir Molaei, Hamid D. Taghirad

この論文をやさしく読む

ひとことで言うと

手術動画の動きと、その動きに含まれる不確かさを使って技能を分類する研究です。ニューラルネットワークの特徴抽出と、木構造のガウス過程分類を組み合わせます。

何に役立つ?

考えられる用途は、手術訓練動画の技能評価を支援することです。学習データ量や計算コストを抑える点も評価対象にしています。

この研究の面白いところ

動作の変化を分類の手掛かりとして使うだけでなく、入力がどれほど不確かかを表す情報としても使います。動きの表現とカーネル設計を組み合わせている点が特徴です。

どこまで分かった?

96.9%はJIGSAWSの被験者独立LOUOでの平均正解率です。被験者内LOSOや別データセットの値と混同できません。原文ではこの数値が未展開の書式マクロで囲まれています。臨床現場全般での評価妥当性までは示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

提案するパイプラインは、representation-flow畳み込みニューラルネットワークと、ダイナミクスおよび不確かさを考慮した木構造のガウス過程分類器を統合する。この枠組みでは、潜在的な運動ダイナミクスを、識別に役立つ表現としても、入力の不確かさの源としても利用し、時間的な変動や異常な動作の遷移に対する頑健性を高める。従来の深層学習手法と比べ、提案手法は必要な学習データが少なく、計算効率も高い。 分類性能をさらに改善するため、手術動画の特徴に埋め込まれた意味、フロー、動的な情報を効果的に捉える、新しい意味を考慮した複合カーネルを導入する。加えて、不確かさを考慮するカーネルを開発し、複合カーネルの枠組みの頑健性と実用上の適用可能性を強化する。 提案手法を、JIGSAWSとCataract-LMM(前嚢切開)という2つのベンチマークデータセットで評価する。実験は両データセットで高い性能を示し、JIGSAWSのLOSOおよびLOUO評価も含む。特に、被験者を独立に分けるJIGSAWSのLOUOでは平均正解率96.9%に達した。被験者内のLOSO結果は先行研究との比較のために報告しており、計算コストを大きく減らしながら競争力のある正解率を得た。全体として、提案パイプラインは、動画に基づく手術技能評価のための効率的で正確な枠組みを提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The proposed pipeline integrates a representation-flow convolutional neural network with a dynamics- and uncertainty-aware tree-based Gaussian Process classifier. In this framework, latent motion dynamics are exploited both as discriminative representations and as a source of input uncertainty, enhancing robustness against temporal variations and abnormal motion transitions. Compared with conventional deep learning approaches, the proposed strategy requires less training data and offers improved computational efficiency. To further improve classification performance, we introduce novel semantic-aware compound kernels that effectively capture semantic, flow, and dynamic information embedded in surgical video features. In addition, uncertainty-aware kernels are developed to strengthen the robustness and practical applicability of the compound kernel framework. The proposed method is evaluated on two benchmark datasets, namely the JIGSAWS and the Cataract-LMM (Capsulorhexis) datasets. Experimental results demonstrate strong performance across both datasets, including the LOSO and LOUO evaluation protocols on JIGSAWS, including the subject-independent LOUO protocol on JIGSAWS, on which the framework attains a mean accuracy of \ph{96.9}\%; results under the within-subject LOSO protocol are reported for comparability with prior work, achieving competitive accuracy while substantially reducing computational cost. Overall, the proposed pipeline provides an efficient and accurate framework for video-based surgical skill assessment.

著者のコメント

4 figures, 17 tables, 31 pages. It is Under Review in scientific reports Journal

arXiv ID: 2609.24619 / 要約の誤りについて