記録済みロボット実演を再描画する視覚領域ランダム化
ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning
この論文をやさしく読む
ひとことで言うと
収録済みのロボット実演動画の背景や物体色を生成で変え、操作データを撮り直さずに見た目の多様性を増やします。
何に役立つ?
模倣学習が収録時の色や背景だけに適応してしまう問題を緩和するためのデータ拡張です。
この研究の面白いところ
映像と指示は変える一方、動作列と身体状態のラベルはそのまま使います。2つの実機で、色を変えた物体への成功率が0%から42.9%と47.5%へ上がっています。
どこまで分かった?
元の外観では性能を保ったと報告しますが、変更した物体での成功率は依然半分未満です。見た目の多様化の効果であり、物体の力学的性質まで変更した実演の実証ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
模倣学習されたロボット方策は、学習用実演に含まれる視覚条件に過適合することが多い。そのため、物体の色や背景の見た目の変化だけで性能が大幅に低下することがある。一般的な対策は、新しい視覚文脈ごとに追加実演を収集することだが、対象とする外観条件のたびにロボット、管理された環境、人間の操作へ繰り返しアクセスする必要があり、資源を大量に消費する。本研究ではReShootを導入する。これは、以前に記録した実演を外観変更後の環境で再レンダリングして視覚的多様性を合成し、負担をデータ収集から生成へ移す枠組みである。視覚言語モデルがシーンをキャプション化し、背景、物体色、材質など狙った属性を編集し、エッジ条件付き動画生成器が2台のカメラ映像を整合するよう再レンダリングする。その後、指示も更新する。行動系列と固有受容感覚軌跡はラベルを付け直さずそのままコピーするため、生成された各エピソードは記録済みの行動ラベルと固有受容感覚ラベルを保持する。 LIBEROでは、記録実演と再レンダリング実演を同量混ぜて学習した方策は、記録実演だけでの学習と同等の性能を示した(96.5%対96.9%)。さらにLIBERO-Plusでは、混合データセットによってシーン摂動への頑健性が向上した(85.5%対82.3%)。2種類の実機ロボットで、事前収集した43件および100件の実演にReShootを適用すると、色を変えた物体での成功率はそれぞれ0.0%から42.9%、47.5%へ上昇し、元の記録時の外観に対する性能は維持された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each novel visual context; however, this approach is resource-intensive, requiring repeated access to a robot, a controlled environment, and human operation for every appearance condition to be covered. We introduce ReShoot, a framework that synthesizes visual diversity by re-rendering previously recorded demonstrations under altered appearances, thereby shifting the burden from data collection to generation. A vision-language model captions the scene, edits a targeted attribute (e.g., background, object color, or material), and an edge-conditioned video generator re-renders both camera views to match. The instruction is updated accordingly. The action sequence and proprioceptive trajectory are copied verbatim without relabeling, so each generated episode retains the recorded action and proprioceptive labels. On LIBERO, a policy trained on an equal mixture of recorded and re-rendered demonstrations matches the performance of recorded-only training (96.5% vs. 96.9%). Moreover, the mixed training set improves robustness to scene perturbations on LIBERO-Plus (85.5% vs. 82.3%). Across two physical robotic platforms, deploying ReShoot with 43 and 100 pre-collected demonstrations increased the success rate on recolored objects from 0.0% to 42.9% and 47.5%, respectively, while maintaining performance under the original recorded appearance.
著者のコメント
Preprint
arXiv ID: 2609.19661 / 要約の誤りについて