arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

ロボットの指令を実機で調整するシミュレーション移行法

Online Sim-to-Real Adaptation via Closed-Loop System Modeling

Yuhao Huang, Samuel A. Moore, and Boyuan Chen

この論文をやさしく読む

ひとことで言うと

学習済みロボット制御器を変えず、実機での観測に合わせて目標指令を調整し、追従誤差を減らす方法です。

何に役立つ?

シミュレーションで作った制御器を実機へ移す際、方策全体の再学習をせずに追従を改善する選択肢になります。

この研究の面白いところ

ロボットと制御方策を一体の閉ループ系として学び、物理系を丸ごと同定せずに指令側を最適化します。

どこまで分かった?

要旨で報告される評価は二足歩行の速度追従と移動しながらの物体操作のシミュレーション・実機試験です。改善量の具体的な数値は要旨にありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

シミュレーションから実機への移行は大きく進歩したものの、動力学の残る不一致によって、実機上で安定して機能する制御器でも目標への追従精度が落ちる場合がある。この誤差を直すには通常、システムの動力学の同定、制御方策の適応、または追加学習や微調整のためのシミュレーションへの差し戻しが必要で、いずれも多くのデータと計算を要し得る。本研究は、既存の制御器へ与える目標指令を代わりに調整する枠組みOSRAMを提案する。OSRAMは、配備したロボットとその方策を一つの閉ループ動力学系として扱い、追従の観測結果から、課題レベルでの指令と応答の関係を直接学習する。閉ループ動力学モデルは、シミュレーション中にランダム化した動力学を用いてメタ学習し、配備後は限られた実機との相互作用で素早く微調整する。次に、適応したモデルを使って将来の目標指令を最適化し、元の制御方策は変更しない。 二足歩行の速度追従と、移動しながらの物体操作について、シミュレーションと実機でOSRAMを評価した。未知の動力学の下で閉ループのモデル化が予測と追従の精度を改善し、オンラインで目標を適応させることで、異なる制御目的とハードウェア構成において、シミュレーションから実機への移行後に残る追従誤差が減った。これらの結果は、方策を微調整したり物理的な動力学全体を同定したりする方法に対し、ロボットと方策を合わせた閉ループの振る舞いを調整する実用的な代替法を示している。詳しい情報は著者らのOSRAMのウェブサイトにある。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Sim-to-real transfer has made substantial progress, but can still produce controllers that remain stable and functional on hardware while suffering from degraded tracking accuracy due to residual dynamics mismatch. Correcting these errors typically requires identifying the underlying system dynamics, adapting the control policy, or returning to simulation for additional training and finetuning, all of which can require substantial data and computation. We propose OSRAM (Online Sim-to-Real Adaptation via Closed-Loop System Modeling), a framework that instead adapts the reference commands provided to an existing controller. OSRAM treats the deployed robot and its policy as a unified closed-loop dynamical system and learns its task-level command-response behavior directly from tracking observations. A closed-loop dynamics model is meta-trained across randomized dynamics in simulation and rapidly finetuned after deployment using limited real-world interaction. The adapted model is then used to optimize future reference commands while leaving the underlying control policy unchanged. We evaluate OSRAM on bipedal velocity tracking and loco-manipulation in simulation and on hardware. Results show that closed-loop modeling improves prediction and tracking accuracy under unseen dynamics, while online reference adaptation reduces residual sim-to-real tracking errors across different control objectives and hardware configurations. These results demonstrate that adapting the behavior of the robot-policy closed loop provides a practical alternative to finetuning the policy or identifying the full physical dynamics for sim-to-real transfer. More information can be found at http://generalroboticslab.com/OSRAM.

arXiv ID: 2609.28878 / 要約の誤りについて