指先の現在と目標の力を画像に示すロボット操作手法
VisForce: Visual Grounding of Current and Desired Forces for Goal-Conditioned Dexterous Manipulation
この論文をやさしく読む
ひとことで言うと
指先で加えている力と目標の力を画像上の指先に重ね、ロボットの動作に反映させる手法。
何に役立つ?
壊れやすい物をつかむなど、力の調整を伴うロボット操作の設計に役立つ可能性がある。
この研究の面白いところ
実機で卵や歯磨き粉のチューブを持ち上げ、さらに三つの複数段階課題で最終成功率を測った。
どこまで分かった?
報告された成功率は実験に使ったロボットと課題での値であり、未評価の物体や環境への汎化は要旨から分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚・言語・行動(VLA)モデルは汎用的なロボット操作方策として登場している。しかし器用な手の操作では、接触力を別の状態や力専用の表現として与えることが多く、力と画像上の位置の対応を明示しにくい。本研究は、現在の力と目標の力を対応する指先位置に視覚的に結び付けるVisForceを提案する。現在の手首カメラ画像と課題ごとの目標画像に力を示す視覚的な手掛かりを描き込み、目標に条件付けたクロスアテンションで二つの表現を組み合わせて、力を考慮した行動を生成する。力を指定する把持課題と三つの複数段階操作課題について、多指ハンドRH56F1を備えた実機ロボットUR10で評価した。把持実験では、目標の力を上げると把持力も一貫して変化し、卵と歯磨き粉のチューブの把持・持ち上げ成功率はそれぞれ70%、80%だった。カップ挿入とボトルからの注ぎ、トングを使ったパンの移動、滑りを調整するペグ挿入の最終成功率は、それぞれ70%、55%、40%だった。これらの結果は、指先位置に合わせた力の視覚表現を、VLAに基づく器用な手の操作で力を考慮する条件として利用できることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-Language-Action (VLA) models have emerged as general-purpose robotic manipulation policies. However, in dexterous hand manipulation, contact forces are typically provided as separate states or force-specific representations, making it difficult to explicitly represent the spatial correspondence between force and their corresponding visual locations. In this work, we propose VisForce, which visually grounds the current and desired forces at their corresponding fingertip locations. VisForce renders current and desired visual force cues on the current wrist image and a task-specific goal image, and combines the two representations through goal-conditioned cross-attention to generate force-aware actions. We evaluate VisForce using a real UR10 robot equipped with an RH56F1 dexterous hand through force-conditioned grasping and three multi-stage manipulation tasks. In force-conditioned grasping experiments, VisForce exhibited a consistent grip-force response as the desired force increased, and achieved grasp-and-lift success rates of 70% and 80% for an egg and a toothpaste tube, respectively. It further achieved final success rates of 70%, 55%, and 40% on cup insertion/bottle pouring, tong-assisted bread transfer, and slip-modulated peg-in-hole, respectively. These results show that fingertip-aligned visual force representations can be effectively used for force-aware conditioning in VLA-based dexterous hand manipulation.
arXiv ID: 2609.25785 / 要約の誤りについて