arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

モーターのトルク差を使った把持動作の実機移行

Simple Torque-Observation Alignment for Zero-Shot Sim-to-Real Grasping with a Direct-Drive Gripper

Doyoung Kim, Edgar Lee, Hyeonsun Park, Chunghyeon Lee, Chihyun Han, Uisu Hwang, and Seokhwan Jeong

この論文をやさしく読む

ひとことで言うと

シミュレーションと実機のトルク観測を合わせ、学習した把持動作を実機に移します。

何に役立つ?

ダイレクトドライブ式グリッパーの方策を、実機データで追加学習せずに使う際の観測調整に役立つ可能性があります。

この研究の面白いところ

トルクの尺度を較正し、絶対値ではなく時間差分を使って一定の偏りを取り除きます。

どこまで分かった?

実機での検証は分布内の9物体に対する比較で、提案法の成功率は100%でした。分布外物体の結果は要旨に記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

強化学習でトルクを観測に使うのは、シミュレーションと実測のトルクで尺度、ずれ、雑音が異なるため難しい。本研究では、モーター電流がモーターの種類に固有のトルク定数Kτを介して関節トルクに線形に対応するダイレクトドライブ(DD)式アクチュエーターのロボット向けに、簡単なトルク観測の整合法を提案する。第1に、動力計による較正でKτ*を求め、シミュレーションと実機のトルクの尺度差を補正する。第2に、領域ごとの偏りを持つトルクτ(t)をそのまま使わず、両領域で差分Δτ(t)=τ(t)−τ(t−1)を観測として使い、一定のずれを除く。第3に、動力計の測定データから得たガウス雑音を学習時に加える。 提案法の検証として、教師・生徒型の把持方策を完全にシミュレーションで学習させ、知識を移した生徒方策を多指のDDグリッパーに導入する。導入した方策は関節位置とトルクの差だけを使って、自己受容的に物体を把持する。分布内の9個の物体について、別の整合法と比較するアブレーション実験を行った。提案法の把持成功率は100%だった。この結果は、実機のトルク観測の不一致に対し、DDグリッパーでの方策の事前実機学習なしの移行をより頑健にすることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Torque observations in reinforcement learning remain challenging because simulated and measured torque differ in scale, offset, and noise. In this paper, we propose a simple torque observation alignment method for robots with direct-drive (DD) actuators, in which motor current maps linearly to joint torque through a motor-type-specific torque constant K_tau. First, dynamometer calibration identifies K_tau* and corrects the scale mismatch between simulated and real torque. Second, the method uses delta_tau(t) = tau(t) - tau(t-1) as the observation in both domains to eliminate the constant offset instead of using the direct torque tau(t), which carries a domain-dependent bias. Third, Gaussian noise obtained from the dynamometer measurement data is injected during the learning process. To validate the proposed method, we train a teacher-student grasping policy entirely in simulation and deploy the distilled student on a multifingered DD gripper. The deployed policy performs proprioceptive grasping using only joint positions and torque differences. We conduct an ablation study comparing the proposed method with alternative alignment variants on nine in-distribution (ID) objects. The proposed method achieves 100% grasp success. These results demonstrate that the proposed alignment method improves the robustness of zero-shot policy transfer on the DD gripper against real-world torque-observation mismatches.

arXiv ID: 2609.29031 / 要約の誤りについて