arXiv論文メモ
新着一覧
cs.RO / cs.CV · 査読状況未確認

人の触覚データで器用なロボットの動作予測を学ぶ

DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

Yan Qin, Yue Chen, Wenwei Lin, Shujia Liu, Chuqiao Lyu, Kailun Su, Weiyang Jin, Chenze Yu, Ping Luo, Wenbo Ding, Tianxing Chen and Renjing Xu

この論文をやさしく読む

ひとことで言うと

人が物を触って動かすデータを使い、ロボットが次に見る画像と感じる接触を予測するモデルを育てる研究です。

何に役立つ?

高価な実ロボットデータの収集を、人の触覚データで補う方法になります。予測モデルに加え、方策の評価や学習用の合成軌跡の生成にも利用して評価しています。

この研究の面白いところ

人とロボットでセンサー配置と行動表現を対応させ、異なる課題のデータでも転用しています。実ロボット5時間を固定し、人のデータ量だけを増やす実験で効果を確かめています。

どこまで分かった?

共通のセンサー配置と、人の動作をロボットの行動空間へ変換する仕組みを用いた結果です。要旨には予測改善の具体的な割合や方策の成功率はなく、任意のロボットにそのまま移せるとまでは示していません。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

接触を多く伴う器用な操作の予測モデルを学ぶには、密な触覚相互作用が必要だが、実ロボットでこうしたデータを大規模に集めるには費用がかかり、センサーも機体ごとに固有のものとなりがちである。本研究では、大規模に収集可能な人の触覚から学習し、将来のRGB観測と両側の触覚ダイナミクスを同時に予測する、行動条件付き世界モデルDexTouch-WMを提案する。触覚観測と行動空間に互換性を持たせれば、人とロボットの操作には転用可能な接触ダイナミクスが共通している、という洞察に基づく。 人の手と器用なロボットハンドの両方に、共通の検出素子配置を持つ柔軟な圧抵抗アレイを装着する。そして人の動きをロボットの行動空間へ変換することで、人の相互作用を、実ロボットの予測にも使う同じダイナミクスモデルの教師データにする。DexTouch-WMは、手の解剖学的構造を考慮した触覚トークンと、対応付けた行動条件を用いて、事前学習済みの映像エキスパートと軽量な触覚エキスパートを結合する。 人からロボットへのデータ拡大実験では、実ロボットの教師データを5時間に固定したまま、人の相互作用を0時間から100時間へ増やした。人とロボットの課題集合が重なっていないにもかかわらず、学習に使っていないロボット領域のデータにおける視覚、幾何、接触の予測が大幅に改善した。予測に加えて、方策評価のための代替環境として、また実ロボットの方策学習に使う合成軌跡の生成器として、世界モデルを評価する。これにより、大規模に集められる人の相互作用が、器用なロボットの世界モデルを学習するための補完的なデータ拡大の軸になることを示す。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-18 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.

著者のコメント

Accept to IROS 2026 Workshop RoBoWoMo (Lightning Talk)

arXiv ID: 2609.20649 / 要約の誤りについて