必要な場面だけ人が実演するロボット作業計画
TANDEM: Task and Motion Planning with As-Needed Demonstrations for Efficient Vision-Language-Action Model Fine-tuning
この論文をやさしく読む
ひとことで言うと
ロボットができる工程は自動化し、できない工程だけ人に遠隔操作してもらうことで、学習用の実演を効率よく集める仕組み。
何に役立つ?
長い工程のロボット操作について、人の介入時間を抑えながらVLAモデルの追加学習データを集める方法として役立つ可能性がある。論文の評価では5タスクを使い、追加学習後の平均成功率が上がった。
この研究の面白いところ
人の作業を計画器の要求時に呼び出せる操作として表し、人の介入後には場面を再認識して結果を確かめる。さらに計画器の動作を事前学習モデルに合う分布へ寄せる。
どこまで分かった?
要旨で示された評価は、計画領域を超える操作タスク5件に基づく。2.9倍の収集量は代表的な1タスクで、平均成功率60%は各タスク20件の実演で追加学習した条件での結果。ほかのタスクへの一般化は要旨からは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
遠隔操作の担当者は、ロボットがすでに自律実行できる動作の実演にも多くの時間を費やしており、ロボットの基盤モデル向けデータ収集を拡大するうえで制約になっている。タスク・動作計画(TAMP)は多くの動作を自動化できるが、固定された計画領域では長い工程のすべてを扱えない場合がある。本研究では、計画器の能力を超えるタスクの実演を収集するため、TAMPと必要なときだけ行う人の遠隔操作を組み合わせたTANDEMを提案する。 中心となる考え方は、人の支援を要求時に利用できる計画能力として表すことだ。言語指示と視覚観測を受け、事前学習済みの視覚言語モデルを用いて、不足する述語と人が実行する「マジック演算子」を計画領域に追加する。これにより、タスクごとに介入箇所を事前指定せず、自律工程と人が実行する工程を交互に組み込める。人が実行した後は場面を再認識し、意図した効果が生じたことを確認してから自律計画を再開する。また、視覚言語行動(VLA)モデルの追加学習に向け、事前学習軌跡の例を使って計画器が生成した動きを対象モデルの事前学習データ分布に合わせる。 TAMPの計画領域だけでは扱えない長い工程の操作タスク5件で評価した。代表的なタスクでは、人の介入時間が同じ場合、全工程を遠隔操作する方法の2.9倍の実演を収集した。各タスク20件のTANDEM実演で事前学習済みのπ₀.₅-DROIDモデルを追加学習すると、5タスクの平均成功率は0%から60%に上昇した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Human teleoperators spend substantial time demonstrating behaviors that robots can already perform autonomously, limiting the scalability of data collection for robot foundation models. Task and motion planning (TAMP) can automate many of these behaviors, but a fixed planning domain may not support every stage of a long-horizon manipulation task. We present TANDEM (Tamp with As-Needed Demonstrations for Efficient Model fine-tuning), a system that combines TAMP with selective human teleoperation to collect demonstrations for tasks beyond the planner's capabilities. Our key idea is to represent human assistance as an on-demand planning capability. Given a language instruction and visual observation, TANDEM uses pretrained vision-language models to extend the planning domain with missing predicates and human-executed magic operators. This allows the planner to interleave autonomous and human-executed stages without task-specific intervention points. After each human stage, TANDEM re-perceives the scene and checks whether the intended effects hold before resuming autonomous planning. To support fine-tuning vision-language-action (VLA) models, TANDEM also uses example pretraining trajectories to align planner-generated motions with the target model's pretraining distribution. We evaluate TANDEM on five long-horizon manipulation tasks beyond the TAMP domain's capabilities. On a representative long-horizon task, TANDEM collects 2.9x as many demonstrations as full-task teleoperation at the same human intervention time. Fine-tuning a pretrained \pi_{0.5}-DROID model on 20 TANDEM demonstrations per task increases average task success from 0% to 60% across the five tasks.
著者のコメント
Under review. Project page: https://prpl-group.com/tandem/. The first two authors contributed equally
arXiv ID: 2609.28314 / 要約の誤りについて