視覚・触覚・音を同時に記録するロボット操作データ収集
PolyUMI: Accessible Visual-Tactile-Audio Data Collection for Object Inference and Manipulation
この論文をやさしく読む
ひとことで言うと
手持ち機器で視覚、触覚、接触音を集め、同じセンサーをロボットへ移して操作を学ぶ仕組みです。
何に役立つ?
ロボットの物体把持や滑り制御の学習用データを集める用途が考えられます。要旨ではこれらの課題での実験結果を示しています。
この研究の面白いところ
実演とロボット実行でセンサーの配置を揃え、触覚と音が視覚を補う効果を調べています。
どこまで分かった?
要旨には各課題の具体的な性能値や、異なる機器への適用範囲は記されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人間は物を扱う際、視覚、触覚、聴覚、手足の位置感覚を合わせて接触を知り、動作を調整する。ロボットにも同程度の応答性を持たせるには、こうした相補的な感覚信号を保存し利用できる機器が必要である。しかし模倣学習システムの多くは主に視覚と位置感覚で実演を観察するため、目では分かりにくい接触の情報を得にくい。著者らは、視覚・触覚・音を使う実演データの大規模収集とロボットへの導入のため、オープンソースのPolyUMIを提示する。軽量で無線の手持ちグリッパーは、作業台に接続した計算機を必要とせず、手首のカメラ、光学式の触覚、接触音、位置感覚の観測を同期記録する。同じ感覚機能を持つ指をロボットの先端へ移せるため、実演収集と方策実行で感覚を取る配置を保てる。 異なる種類の観測を使うため、センサー間と時間方向の情報をトークン単位で統合し、接触を考慮したロボット動作を予測する多モード方策VisTAも提案する。物体の推定、滑りの制御、接触の多い操作にわたる実験では、触覚と音が視覚だけでは得られない課題関連情報を与え、VisTAは既存の多モード方策と同程度かそれ以上の性能を示した。PolyUMIとVisTAは、複数感覚の実演を集め、視覚を超えて物理的な相互作用を捉える方策を学ぶための利用しやすい一連の仕組みを提供する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Humans typically rely on vision, touch, hearing, and proprioception to perceive contact and adapt their actions during manipulation. Providing robots with comparable responsiveness therefore requires hardware that can retain and use these complementary sensory signals. Most imitation-learning systems, however, observe demonstrations primarily through vision and proprioception, limiting access to contact information that is difficult to infer visually. We present PolyUMI, an open-source platform for scalable visual--tactile--audio demonstration collection and robot deployment. Its lightweight, wireless handheld gripper records synchronized wrist-camera, optical tactile, contact-audio, and proprioceptive observations without requiring a tethered workstation. The same sensing finger can be transferred to the robot end effector, preserving the sensing geometry between demonstration collection and policy execution. To effectively use these heterogeneous observations, we further introduce VisTA, a token-level multimodal policy that integrates information across sensors and time to predict contact-aware robot actions. Experiments spanning object inference, slip control, and contact-rich manipulation show that touch and audio reveal task-relevant information beyond vision and that VisTA is competitive with or outperforms existing multimodal policies. Together, PolyUMI and VisTA provide an accessible pipeline for collecting multimodal demonstrations and learning policies that perceive physical interaction beyond vision. Project Page: https://polyumi-vista.github.io
著者のコメント
9 pages, 10 figures, preprint
arXiv ID: 2609.29760 / 要約の誤りについて