単眼動画から人型ロボットが実行できる動作を直接学ぶ方法
BeyondRetarget: Learning Executable Humanoid Motions Directly from Monocular Video
この論文をやさしく読む
ひとことで言うと
人の単眼動画から、人の骨格動作をいったん復元せずに、人型ロボットが実行する動作を直接生成する方法です。
何に役立つ?
人の動画を人型ロボットの動作の手本として使う際に、身体構造の違いや動作推定の誤差を避ける設計として参考になります。
この研究の面白いところ
ロボット向けの暗黙表現を動画から直接学び、接触を考慮した最適化で時間的な一貫性と物理的な妥当性を高めています。実機とシミュレーションの両方で評価しています。
どこまで分かった?
要旨は精度、頑健性、成功率、遅延の改善を述べていますが、比較対象や改善幅の具体的な数値は記載していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人の動画から実行可能な動作を学ぶことは、人型ロボットが手本となる動きを大量に獲得するための拡張可能な方法となる。しかし従来の手順は、まず人の動作を明示的に表現し、その後、動作のリターゲティングによってロボットの動作へ変換することが多い。この方式は大量の人間のデータを学習に使える一方、人と人型ロボットでは移動の仕組みや関節の自由度構成が大きく異なるため、人の表現を中心に作った動作をロボットが実行しにくい。また、人の動作推定で生じた誤差は必然的にリターゲティング段階へ伝わり、共同の最適化で取り除くことができない。 本研究は、単眼RGB動画をロボットの動作へ直接写す、端から端までの枠組みBeyondRetargetを提案する。明示的な人の動作表現を使わず、視覚観測からロボット向けの暗黙表現を直接学ぶことで、身体構造の違いをまたぐ動作の構造を捉えられる。さらに、ロボットがより実行しやすい動作を生成するため、接触を考慮した動作最適化の仕組みを設計し、時間的な一貫性と物理的なもっともらしさを高める。実験では、生成したロボット動作の精度と頑健性が大きく改善し、シミュレーション環境と実際の人型ロボットの両方で、実行成功率が高く、遅延が小さかった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Learning executable motions from human videos offers a scalable solution for humanoid robots to acquire demonstration motions. However, existing pipelines typically first construct an explicit human motion representation and then convert it into robot motions via motion retargeting. Although such methods can effectively leverage large volumes of existing human data for training, the substantial differences between humans and humanoid robots in locomotion mechanisms and joint degree-of-freedom configurations make motions generated by this human-representation-centric approach difficult to execute on robots. Furthermore, errors introduced during human motion estimation inevitably propagate to the retargeting stage and cannot be eliminated via joint optimization. We propose BeyondRetarget, an end-to-end framework that directly maps monocular RGB videos to robot motions. Discarding the explicit human representation, this framework learns robot-oriented implicit representations directly from visual observations, enabling the model to capture cross-morphology motion structures. To generate motions more suitable for robot execution, we further design a contact-aware motion optimization mechanism to improve temporal consistency and physical plausibility. Experiments show that BeyondRetarget significantly improves the accuracy and robustness of generated robot motions, while achieving higher execution success rates and lower latency in both simulation environments and real humanoid robots.
arXiv ID: 2609.29850 / 要約の誤りについて