視覚・言語・行動モデルの連合学習を比較する実験基盤
An Empirical Study and Open Testbed for Federated Fine-Tuning of Vision-Language-Action Models
この論文をやさしく読む
ひとことで言うと
複数拠点のロボット実演データを共有せずにVLAモデルを追加学習する方法を、シミュレーションと実機で比較した。
何に役立つ?
連合型のロボット学習で、どのパラメータを共有して学ぶかを選び、手法を比較する基盤になる。
この研究の面白いところ
集約アルゴリズムより連合学習の対象範囲が効き、実機では拠点間の違いにより集中型学習に追いつきにくかった。
どこまで分かった?
実機での評価は二つの実験・6課題である。個別化は有用な場合があるが、利用できる全体共通モデルは残らない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
事前学習済みの視覚・言語・行動(VLA)モデルを新しいロボット、環境、課題に適応させるには、各拠点で収集され、しばしば廃棄される実演データが必要である。連合学習は、分散した実演データを共有方策の学習に利用する有望な方法だが、大規模な事前学習済みVLAを適応させられるかは未解決である。また再現可能なベンチマークと再利用できる学習基盤が不足し、既存の結果の比較は難しい。本研究は、現代的な事前学習済みVLA方策3種について、LIBERO操作ベンチマークのシミュレーション40課題と、実機ロボットの二つの実験における実世界の6課題を使い、連合型の追加学習を体系的に調べる。実演データはそれぞれ2拠点と3拠点で収集した。連合学習の対象にするパラメータの範囲、三つの集約アルゴリズム、分布が変わったときの評価など、主要な選択を分析した。得られた知見には、集約アルゴリズムの選択より連合学習する範囲の影響が大きいこと、拠点間の差がシミュレーションで表現されるより強い実機では集中型追加学習に追いつきにくいことが含まれる。一方、異質なデータでも集中型追加学習と同等になり得ること、分布変化に対して集中型以上の頑健性を維持し得ること、各参加者が方策の一部だけを連合学習し残りを局所的に保持する個別化が可能なことも示す。個別化は事前学習が弱い場合に役立つが、利用できる全体共通モデルは残らない。今後の研究と公平な比較のため、モデルや実行環境に依存しない実験基盤decentvlaを公開する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Adapting a pretrained Vision-Language-Action (VLA) model to a new robot, environment, or task requires demonstrations that are collected locally and often discarded. Federated learning is a promising approach to exploiting such distributed demonstrations by learning a shared policy. However, whether it can adapt large pretrained VLAs remains an open question, and a lack of reproducible benchmarks for pretrained VLAs and reusable training frameworks makes existing results difficult to compare. In this paper, we conduct a systematic study of federated fine-tuning of three modern pretrained VLA policies on the 40 simulated tasks of the LIBERO manipulation benchmark, and on six real-world tasks in two real-robot experiments, with demonstrations collected across two and three sites, respectively. Our study analyzes the key choices in this setting, spanning multiple federated parameter scopes, three aggregation algorithms, and evaluation under distribution shift. Based on the study, we derive a series of lessons, including the dominance of the federated scope over the choice of aggregation algorithm and the difficulty of matching centralized fine-tuning on physical robots, where cross-site heterogeneity is stronger than simulation captures. We also highlight opportunities for federated VLA learning, such as the ability to match centralized fine-tuning on heterogeneous data, to remain at least as robust as centralized fine-tuning under distribution shift, and to personalize, with each client federating part of the policy and keeping the rest local, which helps where the policy's pretraining is weak but leaves no usable global model. We open-source \decentvla{}, the model- and runtime-agnostic testbed behind the study, to facilitate future research and fair comparisons in federated VLA learning.
arXiv ID: 2609.22973 / 要約の誤りについて