arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

生成動画から別の場面でも使えるロボット技能を得る

V2-STRep: VLM-Grounded Structured Task Representations for Reusable Robot Skills Acquired from Generated Videos

Yexin Hu, Dongheui Lee

この論文をやさしく読む

ひとことで言うと

生成動画の動きをそのまままねるのではなく、作業の段階や物体との位置関係に分解して、別の場面でも使えるロボット技能にします。

何に役立つ?

毎回新しい動画を作らず、適合する別の指示や配置で技能を再利用することを目指します。六つの実機操作課題で成功率向上と場面間の転用を検証しています。

この研究の面白いところ

画像から得た二次元の手掛かりを深度情報で三次元化し、点・軸・面などの簡潔な幾何表現にします。把持の選択と動作計画を一緒に解き、余った回転の自由度で関節制限に対応します。

どこまで分かった?

転用の結果は、取得に成功した技能と適合する新指示についてのものです。任意の作業への汎化を示したわけではなく、要旨には具体的な成功率の数値はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人間の物体操作動画は、ロボットによる実演なしに技能を獲得するための豊富な動作・相互作用の手掛かりを与える。動画生成モデルは、初期場面の画像と作業指示から実演を合成でき、作業ごとに実演を撮影する必要をなくす。しかし、復元した動作は特定の場面での一つの実現例にすぎず、作業構造、幾何学的関係、制約は明示されないままである。 本研究は、視覚言語モデル(VLM)に基づく構造化作業表現を通じて、生成動画の動作を再利用可能なロボット技能へ変換するゼロショットの枠組みV2-STRepを提案する。この表現は動作の段階、参照対象、作業に関係する制約を指定し、目標を点、点と法線、軸、平面、完全な6次元姿勢という最小限の幾何学構造で記述する。VLMが与える画像空間内の二次元の手掛かりをRGB-D観測で三次元へ持ち上げ、作業の幾何構造と把持姿勢候補を復元する。 幾何構造ごとの規則で動作を新しい場面へ移し、作業制約付き軌道最適化で把持の選択とロボットの動作計画全体を結び付ける。作業要件を保ちながら、残された回転の自由度を使って関節の可動域制限に対応する。実行場面への対応付けと制約を更新すれば、別の動画を生成せずに、新しい適合可能な指示の下で技能を再利用できる。 六つの実世界の操作課題での実験により、ベースラインを上回る実行成功率、獲得に成功した技能の確実な場面間転移、実行時の指示変更への適応を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Human manipulation videos provide rich motion and interaction cues for acquiring robot skills without robot demonstrations. Video generation models synthesize such demonstrations from an initial scene image and task instruction, avoiding the need to record demonstrations for each task. However, the recovered motion captures only one scene-specific realization, leaving task structure, geometric relations, and constraints implicit. We present V2-STRep, a zero-shot framework that converts generated video motion into reusable robot skills through VLM-grounded structured task representations. The representation specifies motion phases, references, and task-relevant constraints, with targets described by minimal geometric structures: points, point-normals, axes, planes, and full 6D poses. VLM-provided 2D image-space cues are lifted into 3D using RGB-D observations to reconstruct task geometry and candidate grasp poses. Geometry-specific rules transfer motion to new scenes, while task-constrained trajectory optimization couples grasp selection with complete robot motion planning. It preserves task requirements while using remaining rotational freedom to accommodate joint limits. Updating deployment grounding and constraints enables reuse under new compatible instructions without generating another video. Experiments on six real-world manipulation tasks demonstrate improved execution success over baselines, reliable cross-scene transfer of successfully acquired skills, and adaptation to changed deployment instructions.

arXiv ID: 2609.20582 / 要約の誤りについて