arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

安定性の範囲を守りながらロボット腕の外乱を学習補償

Stability-aware Residual Reinforcement Learning Framework for Robotic Manipulator Disturbance Compensation

Jihong Kim, Joonhyuk Kwon, Hwa Soo Kim, TaeWon Seo, Hyung-Tae Seo

この論文をやさしく読む

ひとことで言うと

通常の制御が補いきれないロボット腕の誤差を強化学習で補正しつつ、学習器の出力を安定性解析から決めた範囲に収める方法です。

何に役立つ?

摩擦や振動などで精密な動作が乱れる場面の補償に役立ちます。6自由度の実機で、シミュレーションからそのまま移した場合の改善も報告されています。

この研究の面白いところ

外乱の状況を推定して補償を切り替えるだけでなく、学習方策が任意の出力をしても誤差を所定の範囲に収めるよう、状態依存の出力制限を組み込んでいます。

どこまで分かった?

27.8%と38.0%はそれぞれ異なる評価条件での追従誤差の削減です。ISSの保証は導出したモデルと制約に基づくもので、要旨には保証範囲の具体式や比較対象の詳細は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

従来の制御器と外乱オブザーバ(DOB)はマニピュレータの精密な軌道追従の標準的な手法だが、パラメータの不確実性、非線形摩擦、複合的な外乱に弱い。本研究では、解析的なオブザーバと強化学習(RL)方策を組み合わせた、残差強化学習DOBの枠組みを提案する。決定論的な基準制御は信頼できる領域内で動作し、RL方策はモデルが捉えられない残差を明示的に対象とする。 この補償を外乱に応じたものにするため、推定器ネットワークが観測履歴を特権的な外乱の文脈情報に対応付ける。これにより潜在空間を外乱の状況ごとに整理し、外乱が切り替わる際にも迅速に適応できるようにする。安定性を保証するため、入力状態安定性(ISS)の解析から、RL方策に対する状態依存の行動上限を導出して強制した。これによって、方策がどのような出力をしても、閉ループ系の追従誤差を保証された範囲内に閉じ込めることを証明できる。 6自由度マニピュレータでの実験は、外乱推定と追従の一貫した改善を示した。実機への追加適応を行わないゼロショットのシミュレーションから実環境への移行では追従誤差が27.8%減少し、学習中に観測されなかった基部振動の外乱の下では38.0%減少した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Although conventional controllers and disturbance observers (DOBs) are the standard for precision tracking in manipulators, they suffer from parameter uncertainty, nonlinear friction, and compound disturbances. This study proposes a residual reinforcement learning DOB framework that pairs an analytical observer with an RL policy. The deterministic baseline operates within a reliable region, whereas the RL policy explicitly targets the residuals that the model cannot capture. To make this compensation disturbance-aware, an estimator network aligns the observation history with a privileged disturbance context, organizing the latent space by disturbance regime and enabling rapid adaptation across disturbance transitions. To guarantee stability, we derived and enforced a state-dependent action bound on the RL policy from an input-to-state stability (ISS) analysis such that the closed loop provably confines the tracking error to a certified envelope for arbitrary policy outputs. Experiments on a 6-DOF manipulator demonstrated consistent improvements in disturbance estimation and tracking, including a 27.8% tracking-error reduction on real hardware under zero-shot sim-to-real transfer and a 38.0% reduction under a base-vibration disturbance that was not observed during training.

著者のコメント

14 pages, 9 figures

arXiv ID: 2609.21307 / 要約の誤りについて