arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

人の実演中にロボットの接触を可視化する手法

Touch2Robot: Robot Touch in the Human Demonstration Loop

Shengcheng Luo, Xiaoyang Cheng, Hong Ying, Xiaoying Zhou, Jiaming Jiang, Haoran Guo, Wanlin Li, Ziyuan Jiao, Chenxi Xiao

この論文をやさしく読む

ひとことで言うと

人が実演を記録するとき、ロボットの手ならどこに触れるかを画面で示し、学習に使いやすいデータを集める。

何に役立つ?

ロボットの操作学習用実演を集める際、接触の食い違いを減らし、成功する実演を効率よく収集する方法として参考になる。

この研究の面白いところ

4つの実世界課題で実機再生完了率が37.9%から72.1%となり、成功実演1件当たりの収集時間は58.6秒から18.2秒になった。接触再構成のF1値44.2%も併せて報告する。

どこまで分かった?

結果は要旨に示された4課題での比較である。接触再構成のF1値は44.2%であり、触覚の完全な再現を示すものではない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人による実演は、物体操作の学習データを大規模に集める方法になるが、人の接触をロボットの手へ移すと、不安定になったり実行不可能になったりする。一方、対象のロボットで直接実演を集めればこの食い違いは避けられるが、収集費用は大きく増す。この両立のため、本研究は人が実演中に対象ロボットの手が物体へどのように触れるかを見ながらデータを集められるTouch2Robotを提案する。 人による操作中の手の動き、触覚グローブの測定値、物体の動きを記録する。これを使い、記録された人の接触に整合することを優先しつつ、実演された物体の動きを再現する、物体ごとの強化学習方策を導く。得られた動作を、入力される人の観測と物体の形状からロボットの手の構成を求める単一のリアルタイム変換器へ蒸留する。収集中は、予測したロボットの手の構成を、シミュレーション内で追跡した物体の位置・向きに同期させ、ロボットと物体の接触を再構成する。その接触を可視化し、実演者が以後の操作を対象の手に合わせて調整できるようにする。 実世界の4課題では、視覚情報だけのフィードバックと比べ、実機での再生完了率の平均が37.9%から72.1%に改善し、再生に成功する実演1件当たりの収集時間が58.6秒から18.2秒に短縮した。再構成した対象ロボットの手の接触は、実機の触覚測定に対してF1値44.2%を達成した。また、Touch2Robotの実演で訓練した方策は、視覚フィードバックだけの場合と比べて、後段のDiffusion Policyの性能を29.1パーセントポイント高めた。これらの結果は、人の実演収集にロボットの触覚を取り入れることで、器用な操作データの品質と収集効率の両方を改善できることを示す。プロジェクトページ:https://Touch2Robot.github.io/。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-22 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Human demonstrations offer a scalable way to collect manipulation data, but their contacts may be unstable or infeasible when transferred to a robot hand. Collecting demonstrations directly on the target robot avoids this mismatch but substantially increases the cost of data collection. To address this trade-off, we present Touch2Robot, a framework that lets humans collect demonstrations while seeing how the target robot hand would contact the object. We capture human hand motion, tactile-glove measurements, and object motion during human manipulation. These recordings guide object-specific RL policies to reproduce the demonstrated object motion while favoring contacts consistent with the recorded human touch. We distill the learned behaviors into a unified real-time retargeter that maps incoming human observations and object geometry to robot hand configurations. During collection, the predicted robot configuration is synchronized with the tracked object pose in simulation to reconstruct robot-object contacts, which are visualized to help the demonstrator adapt subsequent interactions to the target hand. Across four real-world tasks, Touch2Robot improves average real-robot replay completion from 37.9% to 72.1% over visual-only feedback, while reducing the collection time per replay-successful demonstration from 58.6s to 18.2s. Reconstructed target-hand contacts achieve 44.2% F1 against real-robot tactile measurements, and policies trained on Touch2Robot demonstrations improve downstream Diffusion Policy performance by 29.1 percentage points over visual-only feedback. These results show that bringing robot touch into the human demonstration loop improves both the quality and efficiency of scalable dexterous data collection. Project webpage: https://Touch2Robot.github.io/.

著者のコメント

12 pages, 13 figures

arXiv ID: 2609.24660 / 要約の誤りについて