arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

疎な時刻指定キーフレームで人型ロボットの全身動作を制御

PLAT: Sparse Timed Keyframe Motion Tracking for Humanoid Control via Privileged Latent Transition Learning

Zepeng Wang, Jiangxing Wang, Chao Ma, Xiaochuan Shi, Zongqing Lu

この論文をやさしく読む

ひとことで言うと

人型ロボットに少数の目標姿勢と到達時刻だけを渡して、安定した全身動作を生成する制御方法。

何に役立つ?

細かな動作を逐一指定せず、人型ロボットに長い手順を計画させる制御器の設計に役立つ可能性がある。

この研究の面白いところ

学習時だけ密な目標列を使い、運用時は疎な命令で動作し、シミュレーションに加えてUnitree G1にも導入した。

どこまで分かった?

要旨は実機導入の成功を述べるが、成功率や評価した実機動作の範囲は具体的に示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人型ロボットの動作追従方策は、フレームごとの密な参照動作に依存し、計画や対話的な動作生成のための高位の動作制御器として使いにくい。本研究は、将来の少数のキーフレームと各到達希望時刻だけを方策に与え、連続する目標へ安定した全身動作で到達させる「疎な時刻指定キーフレーム動作追従」を調べる。そのため、特権的な潜在遷移学習を使う三段階の方策学習枠組みPLATを提案する。学習時には密な目標列を特権的な教師信号として利用する一方、運用時には疎な時刻指定キーフレーム命令だけを必要とし、密な動作追従と疎な目標条件付き制御をつなぐ。まず事前学習済みの密な追従の熟練モデルから頑健な動作の事前知識を得る。次にDAgger方式の模倣学習で特権的な潜在事前分布を学び、行動を直接最適化せず潜在遷移を改良する残差強化学習を行う。幅広いシミュレーション実験で、さまざまな計画期間にわたり正確かつ安定した疎な時刻指定キーフレームの追従を維持し、特に長期の命令で高い性能を示した。Unitree G1人型ロボットでの導入成功も、この方式が疎な全身動作制御に有効で実用的であることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Humanoid motion tracking policies rely on dense frame-by-frame references, limiting their use as high-level motion controllers for planning and interactive motion generation. We study \emph{Sparse Timed Keyframe Motion Tracking}, where a policy receives only sparse future keyframes and their desired arrival times, and must execute stable whole-body motions that reach successive goals. We propose \textbf{PLAT}, a three-stage sparse timed keyframe motion tracking policy learning framework with \textbf{P}rivileged \textbf{LA}tent \textbf{T}ransition learning. PLAT bridges dense motion tracking and sparse goal-conditioned control by exploiting dense goal sequences as privileged supervision during training while requiring only sparse timed keyframe commands at deployment. A pretrained dense tracking expert first provides robust motion priors. A privileged latent prior is then learned through DAgger-style imitation, followed by latent residual reinforcement learning that refines latent transitions instead of directly optimizing actions. Extensive simulation experiments demonstrate that PLAT maintains accurate and stable sparse timed keyframe tracking across varying planning horizons, with particularly strong performance under long-horizon commands. Successful deployment on a Unitree G1 humanoid robot further demonstrates the effectiveness and practicality of PLAT for sparse humanoid motion control.

著者のコメント

9 pages, 3 figures, under review

arXiv ID: 2609.25754 / 要約の誤りについて