実演中の触覚接触を使いロボットの注目位置を学習
TACIT: Tactile Contact Supervision for Spatial Attention in Dexterous Manipulation
この論文をやさしく読む
ひとことで言うと
ロボットが実演中にどこへ触れたかを手掛かりに、画像のどの位置へ注目すべきかを学ばせます。人が追加で画像上の点を指定する必要をなくす方法です。
何に役立つ?
少ない実演で、位置が変わった対象物にも近づける操作方策を学ぶ用途が考えられます。実機ではボール配置とペグ挿入を評価しています。
この研究の面白いところ
全試行で対象物への接近に成功しても、最終的な課題成功率は約67~73%です。接近の改善と、その後の接触操作で残る失敗を分けて示しています。
どこまで分かった?
評価した作業空間と配置領域での結果です。推論時にも触覚入力は必要です。3シードの比較ではボールは実機、ペグはシミュレーションであり、接触前と接触時の教師信号には一貫した優劣がありませんでした。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
少数の実演から訓練した視覚運動方策は、対象物の位置変化に確実に追従せず、実演された軌道を再現するだけの場合がある。明示的な注意機構を備える既存の方法は、一般に人の注釈や視覚モデルから空間的な事前情報を得る。本研究では、遠隔操作による実演で測定した触覚接触を用い、追加の点注釈なしに空間的注意を教師あり学習するTACIT(触覚接触が注意に情報を与える手法)を導入する。 接触に先行するカメラ点群上にガウス分布の教師信号を与え、注意ヘッドを学習する。その出力を集約したものを、視覚・触覚を用いる拡散方策の条件として入力する。この教師信号は訓練時だけ使用し、推論時にも触覚観測自体は入力として残る。主たる実機ベンチマークでは、各課題10回の実演と5つの実演配置領域を用い、TACITはボール配置で66.7%、ペグ挿入で73.3%の成功率を達成した。これに対し、入力を揃えた3次元視覚・触覚融合は10.0%と20.0%、視覚のみのDP3は20.0%と43.3%だった。TACITは各課題30試行のすべてで、12秒以内に手のひらと対象物の距離が150 mmとなる接近領域へ入り、残る失敗はすべて到達後に生じた。 実機でのボール課題とシミュレーションでのペグ課題について3つの訓練シードで比較すると、TACITは入力を揃えた融合手法と、構成を揃えながら明示的な注意の教師信号を除いた対照手法を上回った。これは、分岐の容量だけでなく教師信号自体が貢献していることを支持する。接触前の教師信号と接触時の教師信号の間には、一貫した優劣は見られなかった。これらの結果は、測定した触覚接触が、評価した作業空間内において、少数の実演から接近動作を学ぶための有効な空間的教師信号となることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Visuomotor policies trained from a few demonstrations may reproduce demonstrated trajectories without reliably following changes in object position. Existing approaches with explicit attention typically obtain spatial priors from human annotation or visual models. We introduce TACIT (tactile contact informs attention), which uses measured tactile contacts from teleoperated demonstrations to supervise spatial attention without additional point annotation. Gaussian targets over preceding camera point clouds supervise an attention head whose pooled output conditions a visuotactile diffusion policy. Targets are used only during training; tactile observations remain inputs at inference. In the primary real-robot benchmark, with ten demonstrations per task and five demonstrated placement regions, TACIT achieves 66.7% success on ball placement and 73.3% on peg insertion, compared with 10.0% and 20.0% for input-matched 3D visuotactile fusion and 20.0% and 43.3% for vision-only DP3. TACIT enters the 150 mm palm-to-object approach region within 12 seconds in all 30 trials per task; all remaining failures occur after arrival. Across three training seeds on real ball and simulated peg, TACIT outperforms input-matched fusion and an architecture-matched control without explicit attention supervision, supporting the contribution of supervision beyond branch capacity. Pre-contact and contact-time supervision show no consistent ordering. These results demonstrate that measured tactile contact provides effective spatial supervision for approach behavior from few demonstrations within the evaluated workspace.
arXiv ID: 2609.24507 / 要約の誤りについて