arXiv論文メモ
新着一覧
cs.CV / cs.AI / cs.RO · 査読状況未確認

主カメラの画像を一部隠してロボットの過学習を抑える

MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model

Yuxuan Jiang, Jiaying Huang, Ge Wang, Shenhao Yan, Jiahao Yang, Chengsi Yao, Qi Liu, Qing Zhao, Shuguang Cui, Yiming Zhao, Yatong Han, Zhen Li

この論文をやさしく読む

ひとことで言うと

ロボットが主カメラの見え方や訓練時の動きに頼りすぎないように、学習中に画像の一部を隠し、手首カメラも活用させる方法です。

何に役立つ?

少量の操作データでVLAを追加学習するときの、軌跡への過学習を抑える用途が考えられます。シミュレーションの評価に加えて実機実験も報告されています。

この研究の面白いところ

新しい大規模な仕組みを加える代わりに、主カメラから得る情報を少し減らすことで、手元の細かな視覚特徴を学ばせようとします。

どこまで分かった?

23.2%と16.8%が相対的な改善率なのかパーセントポイント差なのかは、要旨だけでは明確ではありません。実機実験の成功率や試行数、隠す割合の詳細も要旨にはありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚・言語・行動(VLA)モデルは、視覚と言語の理解を実行可能なロボット行動へ統合し、ロボット制御のエンドツーエンド学習を可能にする。しかし、実証的な分析から、既存モデルは限られたデータセットでファインチューニングすると、軌跡へ深刻に過学習することが分かった。 モデルが手首カメラの情報を有効に使うよう導くため、マスキングに基づくファインチューニング戦略MaskVLAを提案する。主カメラの視覚情報の一部をランダムに隠すことで、モデルがより細かな、課題に関連する有効な視覚特徴を自律的に学ぶよう促す。この過程から頑健な方策が生まれ、複雑な操作課題に対応する能力と汎化性能が向上する。 本手法をRoboTwin 2.0で包括的に評価し、π₀およびOpenVLA-OFTと比較して、平均成功率をそれぞれ23.2%と16.8%改善した。さらに、実環境のALOHAロボットでの実験でも、本手法の有効性を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-Language-Action (VLA) models integrate vision-language understanding with executable robot actions, enabling end-to-end learning for robot control. However, our empirical analysis reveals that existing models exhibit severe trajectory overfitting when finetuned on limited datasets. To guide the model in effectively utilizing wrist camera information, we propose MaskVLA, a masking-based fine-tuning strategy. By randomly masking a small portion of the main camera's visual information, the model is guided to autonomously learn more fine-grained, task-relevant, and effective visual features. This process leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance. Our method has been comprehensively evaluated on RoboTwin 2.0, achieving an average success rate improvement of 23.2% and 16.8% compared to $\pi_0$ and OpenVLA-OFT, respectively. Furthermore, experiments on real-world ALOHA robots also demonstrate the effectiveness of our approach.

著者のコメント

8 pages, 7 figures. Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

arXiv ID: 2609.23565 / 要約の誤りについて