豚組織で熟練医の脂肪吸引操作を記録したデータセット
HumynexSurg-1: A Curated Expert Liposuction Dataset
この論文をやさしく読む
ひとことで言うと
熟練医が豚の腹部組織で脂肪吸引を行う際の映像、力情報、発話などを同期してまとめた学習用データです。
何に役立つ?
外から見えにくい皮下での操作について、動きと力、医師の判断を対応させてロボット学習を研究する基盤になります。
この研究の面白いところ
映像だけでなく判断の発話を専用ラベルに整理し、将来のセンサ追加でも形式を保てる収録設計を採用しています。
どこまで分かった?
35.6分・14エピソードの豚組織データによる概念実証です。患者での自律手術の安全性や有効性は示していません。実測6自由度姿勢、検証済み力計測、超音波などの追加収録と、今回の公開内容は区別が必要です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ロボット基盤モデルは大量の実演データから操作を学ぶが、そのデータには外科手術が不足している。780時間のOpen-H外科手術コレクション全体でも、同期した力情報を持つデータセットは一つだけで、美容目的の処置を扱うものはない。脂肪吸引は、器具が皮下で動き、外科医が感触と判断に基づいて操作するため、難しい事例となる。Humynex Roboticsは、この種の処置のため、熟練者による整理されたデータセットを構築している。 最初の公開版HumynexSurg-1は、熟練した脂肪吸引外科医が豚の腹部組織を対象に、判断を一つずつ言葉にしながら処置する様子を記録したものである。吸引圧、手にかかる6軸の力・トルク、上方からのRGB-D動画、側面動画、襟元マイクを同期させている。14エピソード、42,738フレーム、35.6分、356発話を含み、発話の95%は脂肪吸引専用のラベル体系へ変換できる。 収録は、方策に必要な量を中心に組み立てた特許出願中のセンシング計画に従う。現在はモデルで取得しているチャンネルも、データ形式を変えずに将来センサへ置き換えられる。この公開版では、器具の動きを側面動画の器具・手の追跡として記録し、力チャンネルを状態として提供する。資金を確保した収録では、実測したハンドルの6自由度姿勢、検証済みの力チャンネル、脂肪層の超音波画像、触診センシングを追加する。 概念実証として、NVIDIA Isaac GR00T N1.7を、独自コードなしで1回あたり1時間未満でこのデータセットに微調整し、記録されたセッションを学習させた。同じエピソードを使った規模拡大の検討は、さらなる改善がどこから得られるかを示す。新しいセッションを追加するたびに、未見セッションでの誤差が低下した。提供物はデータセット、ラベル体系、品質保証報告書、評価手順である。これらの検討結果は、ここで挙げたセンサを使い、脂肪の各部位で多数の短いセッションを収録する次段階の方向を示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Robot foundation models learn manipulation from large demonstration corpora, but surgery is missing from those corpora: across the 780-hour Open-H surgical collection, one dataset carries synchronized force and none covers an aesthetic procedure. Liposuction is the hard case, because the instrument works under the skin and the surgeon operates by feel and by judgment. Humynex Robotics builds curated expert datasets for this kind of procedure. HumynexSurg-1 is the first release: a master liposuction surgeon performing on porcine abdominal tissue while narrating every decision, recorded with synchronized suction pressure, six-axis hand force/torque, top-down RGB-D video, side video and a lavalier microphone -- 14 episodes, 42,738 frames, 35.6 minutes, 356 utterances of which 95% compile into a liposuction-specific label schema. The capture follows a patent-pending sensing plan organized around the quantities a policy needs, so a channel captured today by a model can be upgraded to a sensor tomorrow without changing the data format. This release captures the instrument motion as a tool-hand track in the side video and provides the force channel as state; the funded capture adds a measured 6-DoF handle pose, a validated force channel, ultrasound imaging of the fat layer, and palpation sensing. As a proof of concept, NVIDIA Isaac GR00T N1.7 fine-tunes on the dataset with no custom code in under an hour per run and learns the recorded sessions; scaling probes on the same episodes show where further gains come from: every new session lowers the error on an unseen session. The dataset, its label schema, its quality-assurance reports and its evaluation protocol are the product; the next capture, many short sessions across fat regions with the sensors named here, is what the probes point to.
著者のコメント
15 pages, 7 figures, 13 tables. Version 0 dataset release
arXiv ID: 2609.23885 / 要約の誤りについて