arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

手先の目標だけで腕と二足歩行を協調させる

Learning Holistic Whole-Body Loco-Manipulation with a Bipedal Mobile Manipulator

Zhongyu Chen, Yuxuan Nai, Qian Chen, Yidong Zhu, Chen Jing, Qihan Wang, Xudong Li, Zhizhan Li, Leixin Chang, Liangjing Yang, Hua Chen

この論文をやさしく読む

ひとことで言うと

手先をどこへ向けるかだけを指定すると、腕を伸ばす、姿勢を変える、足を踏み出す動作を一つの制御器が調整します。

何に役立つ?

遠隔操作や学習済み方策など、異なる上位の指示源から同じ全身制御器を使う用途があります。

この研究の面白いところ

足や基部の動きを別に指示せず、6自由度の手先目標で腕と脚を協調させる点を実機で示しています。

どこまで分かった?

要旨は複数の指令方式での実機動作を報告していますが、成功率、荷重範囲、未知環境での頑健性の数値は記載していません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

二足での移動と操作の協調により、ロボットは移動とマニピュレーションを組み合わせ、腕の通常の作業範囲を超えた物体と相互作用できる。この能力を実現するには、バランスを保ちながら、タスク水準の操作目標を腕と脚の協調運動へ変換する低水準の全身制御器が必要である。本研究では、強化学習で訓練し、6自由度のエンドエフェクタ目標を二足の基部とロボットアームの協調動作へ直接写像する、統一された全身制御器を提示する。学習した制御器は、手先の目標だけを与えられると、基部の速度や足の着地点を明示的に指示しなくても、到達動作、姿勢適応、踏み出しを自律的に協調させる。訓練中には報酬ゲーティング戦略が、手先追従、移動、バランス間のトレードオフを調整する。また、時間的文脈の推定器が、時間窓を使うTransformer符号化、再帰的GRUメモリ、補助的な動力学予測を組み合わせ、観測履歴から動力学に関わる情報を抽出する。実機実験では、同じ制御器が、VR遠隔操作、学習した拡散方策、台本化した軌道からの指令の下で、到達動作、姿勢適応、踏み出しを支えることを示す。これにより、多様な操作タスクに共通の手先インターフェースを提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm. Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction to extract dynamics-relevant information from observation history. Real-robot experiments demonstrate that the same controller supports reaching, postural adaptation, and stepping under commands from VR teleoperation, a learned diffusion policy, and scripted trajectories, providing a common end-effector interface for diverse manipulation tasks.

arXiv ID: 2609.18930 / 要約の誤りについて