arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

姿勢の要所を介して言語計画と人型ロボット制御をつなぐ

KINO: A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation

Sitong Chen, Fatemeh Zargarbashi, Jin Cheng, Tianxu An, Stelian Coros

この論文をやさしく読む

ひとことで言うと

言語モデルが細かい関節運動を直接決める代わりに、動作途中の重要な姿勢を選び、別の制御方策がその間の全身運動を実行します。

何に役立つ?

言語指示から拾う・運ぶ・置く動作を組み立てる、人型ロボットの制御設計に役立ちます。シミュレーションだけでなくG1実機でも評価されています。

この研究の面白いところ

高レベルの計画と低レベルの制御をキーフレームでつなぎ、物体の位置や寸法に合わせて変換します。学習時のフレーム選択の工夫が成功率改善につながっています。

どこまで分かった?

キーフレームは事前定義したライブラリから選びます。44%から92%という数値は疎なVLMキーフレームを使う評価条件であり、要旨だけではこの値を実機全課題の成功率と特定できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人型ロボットによる移動と物体操作には、課題の指示や場面の意味を解釈しながら、協調した全身運動を実行する必要がある。本研究では、視覚言語モデル(VLM)の計画と強化学習(RL)の制御の間の中間表現として、動作のキーフレームを使う階層的な枠組みを提案する。各キーフレームは、ロボット全身の目標姿勢と、必要な場合は物体の姿勢を指定する。言語指示、場面の観測、実行のフィードバックを受けて、VLMは事前定義したライブラリから課題に関連するキーフレームを順次選ぶ。選ばれたキーフレームは、物体の姿勢や寸法を考慮して現在の場面に合わせて変換する。その後、キーフレームを条件とする全身方策が、目標に到達するための関節レベルの動作を生成する。 低レベル方策の学習用に、顕著性に基づくキーフレームのサンプリング戦略を導入する。これにより、VLMの疎なキーフレームを使う場合の一連の課題の成功率が44%から92%に向上する。シミュレーションとUnitree G1人型ロボットで、物体の拾い上げ、運搬、配置を評価する。このシステムは片手・両手の操作をともに成功させ、学習用の参照データを超える配置場所にも一般化する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Humanoid loco-manipulation requires robots to interpret task instructions and scene semantics while executing coordinated whole-body motions. We propose a hierarchical framework that uses motion keyframes as an intermediate representation between Vision-Language Model (VLM) planning and Reinforcement Learning (RL) control. Each keyframe specifies a target whole-body robot pose and, when applicable, an object pose. Given a language instruction, scene observations, and execution feedback, the VLM selects successive task-relevant keyframes from a predefined library. The selected keyframes are retargeted to the current scene to account for object poses and dimensions. A keyframe-conditioned whole-body policy then generates joint-level actions to reach these goals. We introduce a saliency-based keyframe sampling strategy for low-level policy training that improves end-to-end task success rate from 44% to 92% when using sparse VLM keyframes. We evaluate our framework on object pickup, transport, and placement tasks in simulation and on a Unitree G1 humanoid. The system successfully performs both one- and two-handed manipulation and generalises to placement locations beyond the training reference data.

arXiv ID: 2609.18869 / 要約の誤りについて