一つの動作例から人型ロボットの物体運搬距離を調整
Learning Distance-Conditioned Object Transport for Humanoid Loco-Manipulation from a Single Motion Clip
この論文をやさしく読む
ひとことで言うと
一度だけ示された運搬動作から、指定された別の距離で物を止められるよう、人型ロボットを学習させる方法です。
何に役立つ?
考えられる用途は、距離ごとの実演を用意せずに物体運搬の指示を変えられる制御です。抱えて運ぶ、押す、引くなど四つのモードで調べています。
この研究の面白いところ
途中の場所を通れることと、そこで作業を終えられることを分けて考えています。元の動作の終了部分を途中の状態に移し、実際に達成した配置で教師データを作り直します。
どこまで分かった?
正規化距離MAEの0.15と0.28、シミュレータでの強化学習の改善、実機の距離調整は区別して読む必要があります。要旨には実機の誤差や試行数の具体値はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動作追従は、一つのリターゲティング済み動作クリップから人型ロボットの移動と物体操作を再現できるが、固定された参照動作で学んだ方策は、主としてその実演で示された運搬結果を再現する。元の軌道は途中の物体変位も通過するものの、運搬の終了が示されるのは終点だけである。本研究ではこの不一致を、終了と通過のギャップとして捉える。すなわち、中間の変位は、終了動作まで完了した結果ではなく、通過状態として観測される。 本研究では、実演された終了区間を運搬途中の状態へ移す、距離条件付き参照動作再構成(DCRR)を導入する。パラメータを固定した追従教師が閉ループの動力学の下で再構成した参照動作を再生し、保持した軌道に実際に達成した物体配置のラベルを付け直して、参照動作を必要としない方策へ蒸留する。この手順は、元の動作に符号化された相互作用の振る舞いから、距離を条件とする教師データを構築する。 Carry、Kick-Push、Crouch-Push、Dragにおいて、DCRR-BCは指示に応じた運搬を実現し、正規化した距離の平均絶対誤差(MAE)は全体で0.15となる。元の動作のみを使う行動模倣では0.28である。強化学習による微調整は、学習用シミュレータとシミュレータ間の転移の両方で、指示への応答と実行の頑健性をさらに改善する。最後に、実機実験によって、四つすべての相互作用モードで運搬距離を調整できることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Motion tracking can reproduce humanoid loco-manipulation from a single retargeted motion clip, but a policy trained on a fixed reference primarily reproduces its demonstrated transport outcome. Although the source trajectory visits intermediate object displacements, transport termination is demonstrated only at its endpoint. We identify this mismatch as the termination-versus-passage gap: intermediate displacements are observed as passage states rather than termination-complete outcomes. We introduce Distance-Conditioned Reference Recomposition (DCRR), which relocates the demonstrated termination segment to intermediate transport states. A frozen tracking teacher replays the recomposed references under closed-loop dynamics, and the retained trajectories are relabeled by their achieved object placements and distilled into a reference-free policy. This procedure constructs distance-conditioned supervision from the interaction behavior encoded in the source motion. Across Carry, Kick-Push, Crouch-Push, and Drag, DCRR-BC produces command-dependent transport with an overall normalized distance mean absolute error (MAE) of 0.15, compared with 0.28 for source-only behavior cloning. RL fine-tuning further improves the command response and execution robustness in the training simulator and under sim-to-sim transfer. Finally, hardware experiments demonstrate transport-distance modulation across all four interaction modes.
arXiv ID: 2609.21467 / 要約の誤りについて