arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

人の動画から多様なロボット実演を作るKnowDemo

KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos

Zhiyuan Gao, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Schäfer, Michael Beetz

この論文をやさしく読む

ひとことで言うと

人の動きをそのまままねるのではなく、作業に本当に必要な条件を動画から読み取り、別の持ち方や手順でも成功するロボット実演を作る研究です。

何に役立つ?

考えられる用途は、実機で大量に実演を集める負担を減らし、操作方策の学習データを多様にすることです。生成したシミュレーションデータを使ったモデルで、3タスクの実世界への転移を確認しています。

この研究の面白いところ

タスクの必須条件と、動画の人がたまたま選んだ手順を区別します。その知識を計画前の候補選別にも使うことで、多様性と計画の成功率の両方を改善しています。

どこまで分かった?

実世界への転移の報告は3つのタスクに限られ、成功率の具体値は要旨にありません。候補の計画成功と実機でのタスク成功は異なる評価であり、すべての生成動作が実機で成功したという意味ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボットの操作方策の学習には通常、大量の実演データが必要であり、実機で収集するには費用がかかる。近年の手法は、人の動画から復元した動きを調整し、得られた軌道をシミュレーションで検証することで、ロボットの実演を生成する。しかし、動きの参照例の調整を中心とする方法は、実演された接触戦略や小タスクの順序を引き継ぐため、行動の多様性を制限し得る。また、タスクの要件や場面内の関係の理解が不十分だと、無効な候補を作ってしまい、実演生成の効率が下がる。 これらの限界に対処するため、人の動画から得た構造化された操作知識を使い、対象の作業空間に向けて多様なロボット実演を生成する枠組みKnowDemoを提案する。タスク自体の要件と、個々の実演に固有の選択を分けるため、視覚言語モデル(VLM)に基づく知識抽出・推論モジュールを開発する。このモジュールは、物体と行動の記述を、推定したタスク条件、参照実演、および許容される実行の変化と結び付ける。 この知識を実行可能な実演へ変換するため、記述を対象場面の実体や形状に対応付け、動作計画とシミュレーションに先立って候補の生成と選別を導く。得られる実演は、異なる接触戦略と妥当な小タスク順序を通じて複数の行動様式を示し、構造化された実行ラベルを伴う。実験では、参照のみを用いる構成より多くの検証済み実行様式が得られ、タスクに基づく把持サンプリングにより候補の計画成功率が向上した。生成データが方策学習に使えることを検証するため、事前学習済みπ₀.₅モデルをシミュレーションデータで微調整し、3つのタスクでシミュレーションから実世界への転移を達成した。プロジェクトページは https://zhiyuan-gao.github.io/knowdemo/ である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Learning robot manipulation policies typically requires substantial demonstration data, which are costly to collect on real robots. Recent methods generate robot demonstrations from human videos by adapting recovered motion and validating the resulting trajectories in simulation. However, methods centered on motion-reference adaptation can limit behavioral diversity by retaining the demonstrated contact strategies and subtask orders, while insufficient understanding of task requirements and scene relations can reduce demonstration generation efficiency by generating invalid candidates. To address these limitations, we propose KnowDemo, a framework that uses structured manipulation knowledge from human videos to generate diverse robot demonstrations for a target workspace. To distinguish task requirements from demonstration-specific choices, we develop a knowledge extraction and reasoning module based on a vision-language model (VLM) that associates object and action descriptions with inferred task conditions, demonstration references, and permissible execution variations. To translate this knowledge into executable demonstrations, we resolve the descriptions against target-scene entities and geometry to guide candidate generation and screening before motion planning and simulation. The resulting demonstrations exhibit multimodal behavior through alternative contact strategies and valid subtask orders, with structured execution labels. Experiments demonstrate additional verified execution modes beyond a reference-only configuration and improved candidate planning success through task-guided grasp sampling. To validate the generated data for policy learning, we fine-tune the pretrained $\pi_{0.5}$ model on simulation data, achieving sim-to-real transfer across three tasks. Project page: https://zhiyuan-gao.github.io/knowdemo/

arXiv ID: 2609.21229 / 要約の誤りについて