触覚センサーの配置を使ってロボットの感覚表現を事前学習
Tactile-JEPA: Topology-Aware Self-Supervised Representation Learning for Distributed Tactile Sensors
この論文をやさしく読む
ひとことで言うと
ロボットの表面に不規則に並ぶ触覚センサーについて、一部の情報を隠して予測することで、使いやすい感覚表現を学ぶ方法です。
何に役立つ?
力の推定、手の中にある物体の向きの推定、操作方策の学習などへの利用が示されています。評価では従来手法より力推定誤差6.3%、姿勢推定誤差20.8%の削減を報告しています。
この研究の面白いところ
触覚を画像と同じ格子として扱わず、センサーの接続グラフを使います。細かな接触と表面全体の状態を学ぶため、二つの尺度で隠す設計です。
どこまで分かった?
改善は三つのデータセットと記載されたセンサー・ロボット構成での評価です。要旨には誤差の絶対値や、すべての下流課題の個別成績はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
触覚は、接触が多い器用な操作を行うロボットにとって、特に視界が遮られる状況で不可欠な感覚である。ロボット学習では事前学習済み画像エンコーダーの使用が一般的だが、触覚エンコーダーは依然として、生のノイズの多い信号から最初から学習されることが多く、それが表現力を制限する可能性がある。既存の自己教師あり学習(SSL)は主に視覚方式の触覚センサーに焦点を当てており、分散型の電子皮膚は十分に扱われていない。しかし、これらのセンサーでは、覆う表面上の感知素子が疎かつ不規則に並ぶという独特の性質があり、画像向けSSL手法の直接的な再利用は最適ではない。 本研究では、触覚センサーの空間配置を利用して接続構造を考慮した表現を学ぶ、効率的な自己教師あり事前学習法Tactile-JEPAを提示する。具体的には、センサーの接続グラフを用いて空間的なマスクを決め、隠されていない残りの感知素子から、隠された素子の埋め込み表現を予測するよう学習する。分析により、有効な触覚表現には、局所的な接触の詳細と触覚面全体の状態の両方を捉える必要があることを示し、二つの尺度のマスキングによってこれを実現する。 磁気式とピエゾ抵抗式のセンサー、異なるロボットの身体構成、単一センサーと対になったセンサーの構成を含む三つの多様なデータセットで、Tactile-JEPAは従来の最先端手法に比べ、力推定誤差を6.3%、手内での姿勢推定誤差を20.8%減らし、方策学習を含む他の下流用途でも一貫した改善を得る。全体として、触覚を用いる利点がエンコーダーの事前学習の品質に大きく依存することを結果は示しており、Tactile-JEPAはこの問題に直接取り組む。コードはhttps://github.com/E-Kovtun/tactileで公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Tactile sensing is an essential modality for robots performing contact-rich, dexterous manipulation, particularly under visual occlusion. While pre-trained image encoders are standard in robot learning pipelines, tactile encoders are still commonly trained from scratch from raw, noisy signals, which might limit their expressivity. Existing self-supervised learning (SSL) approaches focus predominantly on vision-based tactile sensors, leaving distributed electronic skins largely unaddressed. These sensors, however, have a distinctive property: their sensing elements are sparse and irregularly arranged over the surface they cover, which makes direct reuse of visual SSL methods suboptimal. We present Tactile-JEPA, an efficient self-supervised pre-training method that uses the spatial arrangement of tactile sensors to learn topology-aware representations. Specifically, it is trained to predict the embeddings of masked sensing elements from the unmasked remainder, using the sensor connectivity graph to guide spatial masking. Our analysis shows that effective tactile representations require capturing both local contact details and the global state of the tactile surface, which we achieve through dual-scale masking. Across three diverse datasets spanning magnetic and piezoresistive sensors, different robot embodiments, and single- and paired-sensor configurations, Tactile-JEPA reduces force estimation error by 6.3% and in-hand orientation error by 20.8% over the prior state-of-the-art, with consistent gains in other downstream applications, including policy learning. Overall, our results demonstrate that the benefit of tactile sensing depends critically on the quality of encoder pre-training, a problem which Tactile-JEPA addresses directly. Code is available at https://github.com/E-Kovtun/tactile.
arXiv ID: 2609.24385 / 要約の誤りについて