物理的な動作制約を守るロボット手取り教示
PAKT: Physically-Aligned Kinesthetic Teaching for Reinforcement Learning
この論文をやさしく読む
ひとことで言うと
人がロボットを直接動かして強化学習用の実演を集めるとき、実行時にも守れる運動制約の範囲に軌道を収める方法を提案した。
何に役立つ?
接触を伴う工業用ロボット操作で、実行可能な実演を集め、教示や介入の負担を減らす方法を考える材料になる。
この研究の面白いところ
手取り教示で集める軌道に、方策実行時と同じ運動学的制限を適用した。四つのベンチマークで、比較対象に対しサイクル時間23~48%、介入回数62~86%の減少を報告した。
どこまで分かった?
要旨の結果は四つの挿入・組み立てベンチマークで報告された試行に基づく。冒頭で挙げた99%超の成功率など、産業上の全要件を達成したとは記載されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
実世界の強化学習システムは、マイクロメートル単位の精度、99%を超える成功率、人間並みの作業サイクル時間が必要な、接触の多い工業用操作に依然として苦戦する。オフポリシーのアルゴリズムは実演や人の介入を利用して性能を高められるが、物理的なシステムと方策の制約を守りながら、直感的に指導データを集めるインターフェースが不足している。そこで、強化学習向けの手取り教示の枠組みPAKTを提案する。PAKTは遠隔操作ではなく、産業界で広く使われる、操作者がロボットを直接動かす教示に基づく。ただし、この方法では操作者が、ロボットや方策には物理的に再現できない速度、加速度、加加速度などの軌道で動かしてしまう可能性がある。PAKTでは、操作者が加えた力を動作に変換するアドミッタンス制御を通じてロボットを誘導する。後段の参照軌道生成器は、方策実行時と同じ運動学的制限を適用し、収集する軌道をその範囲に収める。 この教示インターフェースを支えるため、PAKTには、低頻度の強化学習行動を高頻度のトルク指令へ変換する高性能な制御系を加える。参照軌道生成器とその後のインピーダンス制御器で構成され、参照軌道生成器はインピーダンス制御器の追従性能を保ちながら、接触時の扱いを改善し、方策の動作を滑らかにする。データセンターの計算機トレイを含む四つの挿入・工業組み立てのベンチマークで報告された試行全体では、端から端までのシステムが、HIL-SERLを基準にサイクル時間を23~48%、介入の累積回数を62~86%減らした。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Real-world reinforcement learning (RL) systems still struggle with the demands of contact-rich industrial manipulation, including micrometer-level precision, success rates above 99%, and human-level cycle times. Although off-policy algorithms can improve performance by leveraging demonstrations and interventions, a key bottleneck is the lack of an intuitive interface for collecting such guidance while complying with constraints of the physical system and the policy. We propose PAKT, a framework for kinesthetic teaching in RL. As opposed to teleoperation approaches, PAKT relies on kinesthetic guidance, which is widely used in industry. However, a critical weakness of kinesthetic guidance is the possibility for the operator to move the robot along trajectories (e.g., velocities, accelerations, jerk) that the robot and/or policy cannot physically reproduce. Using PAKT, operators guide the robot through admittance control, which maps human-applied forces to motion. The downstream reference generator applies the same kinematic limits used during policy execution, keeping the collected trajectories within these limits. To support this teaching interface with an appropriate execution layer, PAKT adds a high-performance control stack that maps low-frequency RL actions to high-frequency torque commands. It consists of a reference generator and subsequent impedance controller, where the reference generator preserves the tracking performance of the impedance controller while improving contact handling and producing smoother policy actions. Across the reported runs on four insertion and industrial assembly benchmarks, including a data center compute tray, the end-to-end system reduces cycle time by 23%-48% and cumulative intervention count by 62%-86% relative to the HIL-SERL baseline. Project website: https://pakt-website.github.io/pakt-website}{https://pakt-website.github.io/pakt-website
著者のコメント
17 pages, 6 figures, 9 tables, Conference on Robot Learning
arXiv ID: 2609.25630 / 要約の誤りについて