視覚と身体感覚で歩きながら扉を閉める人型ロボット
ViLoMan: Learning Visual-Proprioceptive Whole-Body Loco-Manipulation Skills for Humanoid Robots
この論文をやさしく読む
ひとことで言うと
人の動作から学んだ一つの制御方策で、人型ロボットが移動と腕の操作を合わせて扉を閉めます。
何に役立つ?
人の実演をロボットで実行できる動きへ変換し、機体のセンサーだけで動く全身制御を学ぶ方法として役立ちます。
この研究の面白いところ
歩行と操作を別々の中間指令でつなぐのではなく、奥行き画像と身体の状態から関節の動きまで直接決める点が特徴です。
どこまで分かった?
シミュレーションと実機での検証がありますが、要旨で示される評価課題は扉閉めです。成功率などの具体的な数値や、他の操作への汎化結果は記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人型ロボットの移動・物体操作には、移動と物理的な相互作用を途切れなく統合する、状況に応じた全身協調が必要である。近年の進歩にもかかわらず、自律的な移動・物体操作の学習は依然として難しい。多様で、物理的に実行可能なロボットと物体の相互作用データが少なく、機体上の観測から直接、統一的な全身制御を学ぶことも難しいためである。 本研究では、自律的な人型ロボットの移動・物体操作のための、規模拡張可能な枠組みViLoManを提案する。ViLoManはまず、人間と物体の相互作用についての部分的な運動学的実演を、完全で物理的に実行可能なロボットの軌道へ変換する。続いて、教師・生徒型の蒸留の枠組みでこれらの軌道を利用し、自己視点の奥行き観測と固有受容感覚の計測値を、関節レベルの全身動作へ直接写像する統一方策を学習する。実行時には、参照動作も中間的な指令も必要としない。 シミュレーションと実世界の両方で、扉の構成とロボットの初期状態をさまざまに変えた扉閉め課題によりViLoManを評価する。実験結果から、単一の方策によってUnitree G1人型ロボットが、機体上の奥行きセンサーと固有受容感覚だけを使って課題全体を完了できることが示された。また、課題の変化に対して頑健に汎化し、シミュレーションから現実へ効果的に移行する。プロジェクトページ:viloman-anonymous.pages.dev。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Humanoid loco-manipulation requires adaptive whole-body coordination to seamlessly integrate locomotion and physical interaction. Despite recent advances, learning autonomous loco-manipulation remains challenging due to the scarcity of diverse, physically executable robot-object interaction data and the difficulty of learning unified whole-body control directly from onboard observations. We present ViLoMan, a scalable framework for autonomous humanoid loco-manipulation. ViLoMan first transforms partial kinematic demonstrations of human-object interactions into complete, physically executable robot trajectories. It then leverages these trajectories within a teacher-student distillation framework to learn a unified policy that maps egocentric depth observations and proprioceptive measurements directly to joint-level whole-body actions. During deployment, the policy requires neither reference motions nor intermediate commands. We evaluate ViLoMan on door-closing tasks across diverse door configurations and robot initial conditions in both simulation and the real world. Experimental results demonstrate that a single policy enables a Unitree G1 humanoid to complete the full task using only onboard depth sensing and proprioception, while generalizing robustly across task variations and transferring effectively from simulation to reality. Project page: viloman-anonymous.pages.dev.
著者のコメント
10 pages, 7 figures. Project page: https://viloman-anonymous.pages.dev/
arXiv ID: 2609.19340 / 要約の誤りについて