arXiv論文メモ
新着一覧
cs.RO / cs.AI / cs.CV · 査読状況未確認

指先の接触変化を予測し、両手ロボットの操作を改善

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Haoran Yuan, Zekai Wang, Boning Shao, Haoran Lu, Trevor Darrell, Ismini Lourentzou, Wei Zhan

この論文をやさしく読む

ひとことで言うと

指先が今どう触れているかだけでなく、接触が次にどう変わるかを予測してロボットの動作を決めます。

何に役立つ?

複数の指や両手で物を扱う操作の学習に役立ちます。要旨では22自由度の両手構成で、接触の多い6タスクを評価しています。

この研究の面白いところ

触覚特徴を同じに保っても、接触の未来を予測する部分を外すと性能が大きく下がります。触覚入力の追加と触覚世界の予測を比較で分けています。

どこまで分かった?

70.6と38.0は要旨で平均スコアと記された値であり、成功率とは断定できません。4タスクの要素除去と6タスクの主比較、視覚品質のdB差は異なる評価です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

器用な物体操作は、視覚だけでは部分的にしか観測できないことの多い接触ダイナミクスに依存する。近年の世界・行動モデル(WAM)は予測的な動画世界モデルと行動生成を結び付けるが、依然として主に視覚中心であり、接触ダイナミクスを直接モデル化できない。各指先を独立に符号化し、指と姿勢を考慮する触覚圧縮器で特徴を集約して、触覚潜在表現を動画拡散世界モデルへ注入する、視覚・触覚WAMのDexTacWAMを提案する。これにより視覚と触覚の世界を共同でモデル化する。22自由度の両手プラットフォームで、接触の多い器用な操作タスク6種類を評価したところ、全タスクで最高スコアを達成し、平均は70.6で、最も強いベースラインの38.0を上回った。要素除去実験では、向上の要因は単に触覚を条件として与えることではなく、予測する世界状態の一部として接触の変化をモデル化することにあった。同じ触覚特徴と行動エキスパートを保ったまま触覚世界モデル化を除くと、4タスク平均は74.7から26.6へ低下した。事前学習済みの視覚VAEを固定した4時間の触覚エンコーダ適応後、視覚から触覚への継続学習により、触覚の中間学習をせず、タスク当たり約100回の実演で事前学習済み動画モデルを触覚へ拡張できた。その際、視覚予測品質は視覚のみの対応モデルとの差0.5 dB以内に保たれた。圧縮器は融合前の接触再現率の89.4%を維持しつつ、学習を2.26倍、推論を1.29倍高速化した。これらの結果は、事前学習済み動画の事前知識を、データと計算の両面で効率よく、分散した複数指の接触ダイナミクスへ拡張できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a visuo-tactile WAM that encodes each fingertip independently, aggregates the resulting features through a finger- and pose-aware tactile compressor, and injects the tactile latent into a video diffusion world model for joint visuo-tactile world modeling. Across six contact-rich dexterous manipulation tasks on a 22-DoF bimanual platform, DexTacWAM achieves the highest score on every task, averaging 70.6 versus 38.0 for the strongest baseline. Ablations attribute the gain to modeling contact evolution as part of the predicted world state rather than tactile conditioning alone: removing tactile world modeling reduces the four-task mean from 74.7 to 26.6 while keeping the same tactile features and action expert. After four hours of tactile-encoder adaptation with a frozen pretrained vision VAE, our continual vision-to-touch learning extends the pretrained video model to touch using roughly 100 demonstrations per task without tactile midtraining, while retaining visual prediction quality within 0.5 dB of vision-only counterparts. The compressor retains 89.4% of pre-fusion contact recall while enabling 2.26x faster training and 1.29x faster inference. Together, these results show that pretrained video priors can be extended to distributed multi-finger contact dynamics in a data- and compute-efficient manner.

著者のコメント

22 pages. Project website: https://dextacwam.github.io/

arXiv ID: 2609.24976 / 要約の誤りについて