arXiv論文メモ
新着一覧
cs.RO / cs.AI / cs.LG · 査読状況未確認

複数のロボット操作を一つの方策へ蒸留する

Distillation for Efficient Multitask Manipulation Policies via Conditional Flow Matching

Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada

この論文をやさしく読む

ひとことで言うと

個別のロボット操作を学んだ複数の専門モデルから、一つの共通方策へ知識を移す方法。

何に役立つ?

ロボット操作の課題ごとに模型を増やさず、複数課題を一つの模型で扱う設計に役立つ。

この研究の面白いところ

CFM専門模型の速度場を蒸留し、元の実演データへの忠実さも学習目的に残した。

どこまで分かった?

性能向上はRLBenchの実験で示された。実機のロボットでの評価は要旨にない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成モデルの進歩は近年、ロボットの方策学習にも広く使われている。特に、専門家の実演で学習した条件付きフローマッチング(CFM)は、ロボット操作のベンチマークで従来法を上回ると示されている。先行研究は主に単一課題に注目してきたが、課題ごとに独立した模型を学習するのは計算費用が高いため、本研究は複数課題を扱う。一方、複数課題の実演データを単純に連結して学習すると、増えた複雑さに対応するため模型を大きくする必要があるか、性能が下がる。本研究は、単一課題のCFM専門模型が学んだ速度場を移し、共有の複数課題方策へ知識を蒸留する。実演への忠実さを保つため、この蒸留の信号と元のCFMの学習目的を組み合わせる。RLBenchでの実験では、模型の大きさを固定したまま、単純な学習より複数課題方策の性能が向上した。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-24 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Advances in generative modeling have recently been extensively employed in robotics for policy learning. In particular, Conditional Flow Matching (CFM) trained with expert demonstrations has been shown to outperform existing methods on robot manipulation benchmarks. While prior work has mainly focused on single-task settings, we study the problem from a multi-task perspective, as training independent models for each task is computationally expensive. Multi-Task policy learning comes with its own set of challenges, as naively training on a concatenated dataset of demonstrations would either require increased model capacity to accommodate the added complexity or result in drops in performance. We propose to distill knowledge from single-task CFM experts into a shared multi-task policy by transferring their learned velocity fields. We combine this distillation signal with the original CFM objective to retain fidelity to the demonstrations. Experiments on RLBench show that our approach improves multi-task policy performance over naive training while maintaining a fixed model size.

arXiv ID: 2609.28107 / 要約の誤りについて