足の着地点と上半身を同時に制御する人型ロボット手法
STRIDER: Stepping-Enabled Multi-Gait Hierarchical 3D Loco-Manipulation Framework for Humanoid Robots
この論文をやさしく読む
ひとことで言うと
人型ロボットが足を置く場所を細かく決めながら、上半身でも作業できるようにする制御法です。
何に役立つ?
地形に合わせた移動と手先の操作を同時に必要とする作業への利用が考えられます。シミュレーションに加えて実機でも評価しています。
この研究の面白いところ
複数の専門方策の動作をまねるだけでなく、内部の技能表現をそろえる学習を加え、1つの方策に統合します。
どこまで分かった?
評価対象はTianGong Omniで、要旨には追従誤差の具体値や試験地形の範囲はありません。あらゆる地形での動作を保証するものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人型ロボットの移動と操作の統合には、2つの大きな制約がある。連続的な速度指令を使う制御器は個々の足の着地点を正確に調整できず、一方、専用の着地点追従モジュールは全身操作と統合しにくい。さらに、標準的な行動ベースの模倣蒸留は、主に専門方策の動作を伝えるもので、異種の技能に共通する表現を明示的には促さない。 本論文では、これらの隔たりを埋める階層型の複数歩容の枠組みSTRIDERを導入する。地形を考慮した3次元の足運びの論理、敵対的動作事前分布(AMP)に基づく自然な歩行、直交座標による上半身制御を統合する。足運びの専門方策は支持足座標系で実行可能な着地点を選び、障害物との間隔を考慮した遊脚軌道を生成する。 異なる歩行・足運びの専門方策を実行可能な単一の生徒方策に融合するため、教師条件付きの潜在表現整合を加えた蒸留アルゴリズム、潜在蒸留近接方策最適化(LD-PPO)を提案する。オンポリシー強化学習、DAggerによる動作再構成、潜在表現整合を同時に最適化し、専門方策の動作を伝えながら、異種のモード間で技能表現の共有を促す。TianGong Omni人型ロボットでのシミュレーションと実機評価により、LD-PPOは通常の蒸留PPOより着地点と姿勢の追従精度が高いことを示す。実機上のSTRIDERは、着地点とエンドエフェクターを正確に追従させながら、複数歩容での移動・操作を実現する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Humanoid loco-manipulation faces two prominent limitations: controllers using continuous velocity commands cannot precisely regulate individual footholds, while specialized foothold-tracking modules are difficult to integrate with whole-body manipulation. Furthermore, standard action-based imitation distillation primarily transfers expert actions, without explicitly encouraging a shared representation of heterogeneous skills. This paper introduces STRIDER, a hierarchical multi-gait framework to bridge these gaps. The framework integrates terrain-aware 3D stepping logic, Adversarial Motion Priors (AMP)-based natural walking, and Cartesian upper-body control: its stepping expert selects feasible footholds in the stance-foot frame and generates clearance-aware swing trajectories. To fuse distinct walking and stepping experts into one executable student policy, we propose Latent Distillation Proximal Policy Optimization (LD-PPO), a distillation algorithm augmented with teacher-conditioned latent alignment. By jointly optimizing on-policy reinforcement learning, DAgger-based action reconstruction, and latent alignment, LD-PPO transfers expert actions while encouraging a shared skill representation across heterogeneous modes. Simulation and real-robot evaluations on the TianGong Omni humanoid show that LD-PPO outperforms vanilla distillation-PPO in foothold-tracking and posture-tracking accuracy. Deployed on hardware, STRIDER realizes multi-gait loco-manipulation with accurate foothold and end-effector tracking.
著者のコメント
9 pages. Submitted to ICRA 2027. Video: https://youtu.be/gf5RWjCZXtA
arXiv ID: 2609.23483 / 要約の誤りについて