内視鏡映像の組織変形から3次元の力を推定する
Object-Centered Reconstruction for Vision-Based 3D Force Estimation
この論文をやさしく読む
ひとことで言うと
内視鏡の左右画像から組織の変形を追い、専用の力センサーを使わずに、かかっている力の大きさと方向を推定する研究です。
何に役立つ?
考えられる用途は、力センサーのない手術ロボットで術者に力の情報を補うことです。現段階ではファントムと摘出ブタ結腸で定量評価し、生体内映像では定性的な実現可能性を示しています。
この研究の面白いところ
カメラではなく組織側を基準とする座標系を使い、視点や組織の向きが変わっても扱いやすい表現にしています。座標表現の変更と追跡方法の変更の効果を分けて比較しています。
どこまで分かった?
0.77 Nと1.30 Nの誤差は、それぞれファントムと摘出ブタ結腸での値です。生体内映像の評価は定性的であり、患者での力推定精度や縫合不全の減少を実証したとは述べていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ロボット支援大腸・直腸手術では、過大な力によって組織が損傷し、吻合部の縫合不全の危険が高まる可能性がある。da Vinci 5には力の検知機能があるものの、それ以前のda Vinciシステムや多くの他の手術ロボットには備わっていない。本研究では、ステレオ内視鏡映像における軟組織の変形から、3次元の相互作用力を推定する視覚ベースの処理系を提示する。物体を中心とした座標系で組織の点群を動的に再構成し、幾何学的な制約を使って組織上の点を追跡したうえで、ニューラルネットワークによって3次元の力ベクトルを予測する。 ゴム手袋のファントム、摘出したブタの結腸、生体内の大腸・直腸手術映像の順に、段階的に処理系を評価する。内視鏡画像中の組織の向きや位置、カメラの視点を変化させた条件で、平均二乗平均平方根誤差(RMSE)はファントムで0.77 N、ブタ結腸で1.30 Nだった。カメラ座標系での表現と比較すると、物体中心の表現は平均RMSEをそれぞれ51.3%と56.7%低減した。また、幾何学的制約を用いた追跡は、CoTrackerと比較してRMSEをそれぞれ19.8%と25.3%低減した。さらに生体内の大腸・直腸手術映像で視覚に基づく力推定の実現可能性を定性的に示した。これは、映像を使ったセンサー不要の力推定を臨床応用へ進めるための一歩となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Excessive force may damage tissue and increase the risk of anastomotic leakage in robotic colorectal surgery. Although the da Vinci 5 provides force sensing, this capability is unavailable on earlier da Vinci systems and many other surgical robotic platforms. In this work, we present a vision-based pipeline for estimating 3D interaction forces from soft-tissue deformation in stereo endoscopic video. We dynamically reconstruct the tissue point cloud in an object-centered coordinate frame, track tissue points with geometric constraints, and predict the 3D force vector with a neural network. We progressively evaluate the pipeline on rubber-glove phantoms, ex vivo porcine colons, and in vivo colorectal surgical video sequences. Under varying tissue orientations and positions within the endoscopic view, as well as different camera viewpoints, the proposed method achieves average root mean square error (RMSEs) of 0.77 N and 1.30 N on the phantom and porcine colon, respectively. Compared with the camera-frame representation, the object-centered representation reduces average RMSE by 51.3% and 56.7%, while geometry-constrained tracking reduces RMSE by 19.8% and 25.3% compared with CoTracker. We further qualitatively demonstrate the feasibility of vision-based force estimation on an in vivo colorectal surgical sequence, as a step toward clinical translation of vision-based, sensorless force estimation.
arXiv ID: 2609.23856 / 要約の誤りについて