未来の接触を予測してロボットの精密操作を改善
PSR: Predictive Sensorimotor Representation Learning for Contact-Rich Manipulation
この論文をやさしく読む
ひとことで言うと
ロボットが現在感じている力に反応するだけでなく、次にどのような接触が起こるかを予測して、精密な動作を選べるようにする研究です。
何に役立つ?
考えられる用途は、物体との接触力を扱う必要がある精密なロボット操作です。要旨では実環境の6課題で全体成功率91.7%を報告しています。
この研究の面白いところ
未来の相互作用の予測を事前学習し、そこで得た表現を動作生成の複数の深さへ渡しています。力の信号を単に入力へ追加することから一歩進め、予測能力を操作に利用する構成です。
どこまで分かった?
成功率と比較改善は、評価した6つの実環境課題についての値です。30.0などの改善値は相対的な割合ではなくパーセントポイント差です。要旨には試行数や課題別の成績は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
接触の多い操作では、視覚情報だけでなく、接触力、ロボットの構成、相互作用の履歴について推論し、方策が精密な動作を生成する必要がある。既存手法は、将来の接触ダイナミクスを能動的に予測するのではなく、力のフィードバックを受動的な条件として用いており、高精度な動作を生成する能力が制限される。 この問題に対処するため、多様な感覚運動信号から予測表現の階層を学習し、それを視覚運動方策の動作ストリームへ統合する、予測的感覚運動表現(PSR)学習の枠組みを導入する。具体的には、事前学習段階でマルチモーダルTransformerを訓練し、将来の相互作用ダイナミクスを共同で予測することで、階層的な予測表現を学ばせる。その後、学習した階層で動作ストリームを拡充し、得られた方策が複数の深さで接触に関する手掛かりを利用できるようにする。 さらにPSRを視覚・言語・動作(VLA)モデル内に実装したPSR-VLAを構成し、実環境の6つの接触の多い操作課題で評価する。実験では、PSR-VLAは全体で91.7%の成功率を達成し、π₀.₅、ForceVLA-π₀.₅、ForceVLA2-π₀.₅を、それぞれ30.0、22.5、19.2パーセントポイント上回った。これらの結果は、力を考慮した接触の多い操作に対するPSRの有効性を示す。課題の動画と安定性試験は https://psr-vla.pages.dev/ で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Contact-rich manipulation requires policies to generate precise actions by reasoning over contact forces, robot configurations, and interaction histories beyond visual observations. Existing methods passively condition on force feedback rather than actively predicting future contact dynamics, limiting their ability to generate high-precision actions. To address this problem, we introduce Predictive Sensorimotor Representation (PSR) learning, a framework that learns a hierarchy of predictive representations from multimodal sensorimotor signals and integrates them into the action stream of a visuomotor policy. Specifically, during a pretraining stage, a multimodal Transformer is trained to learn a hierarchy of predictive representations by jointly forecasting future interaction dynamics. The learned hierarchy subsequently augments the action stream, enabling the resulting policy to exploit contact-relevant cues at multiple depths. We further instantiate PSR within a Vision-Language-Action (VLA) model, resulting in PSR-VLA, and evaluate it on six real-world contact-rich manipulation tasks. Experimental results show that PSR-VLA achieves 91.7% overall success, improving over $\pi_{0.5}$, ForceVLA-$\pi_{0.5}$, and ForceVLA2-$\pi_{0.5}$ by 30.0, 22.5, and 19.2 percentage points, respectively. These results demonstrate the effectiveness of the proposed PSR for force-aware, contact-rich manipulation. Videos of the tasks and stability tests are available at https://psr-vla.pages.dev/.
著者のコメント
7 pages, 5 figures
arXiv ID: 2609.21753 / 要約の誤りについて