arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

ロボット操作で強化学習の専門制御へ切り替える

RouteRLT: Learning When and Which RL Specialist Should Control a Vision-Language-Action Policy

Chongyu Zhu, Jaden Hinds, Hyegang Kim, Juan Sebastian Rojas, Ramy Elmallah, Chi-Guhn Lee

この論文をやさしく読む

ひとことで言うと

ロボット作業の精密な段階だけ、汎用モデルから専門の強化学習方策へ制御を切り替える。

何に役立つ?

把持や挿入など複数の段階を含む作業で、汎用性と局所的な精度を両立させる設計に役立つ。

この研究の面白いところ

段階の境界情報を実行時に与えずに切り替えを学び、シミュレーションとケーブル操作の実機で評価した。

どこまで分かった?

実機で示したのは指定されたケーブル把持・挿入課題と引き継ぎ手順での結果。広範な産業作業への一般化は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚・言語・行動(VLA)モデルは幅広い操作ができるが、コネクター挿入やケーブル取り回しのように接触が多い産業作業では、精密さが重要な段階で苦戦しがちである。事前学習済みVLAを強化学習(RL)で改善すれば、行動模倣を超える課題固有の向上が見込める。しかし、汎用的な振る舞いを保ちつつ、いつ改善が必要で、どの専門方策を動かすかを決める方法は未解決である。RouteRLTは、精密さが重要な一つの段階のために学習したRL専門方策へ、いつ、どの方策へ汎用VLAから制御を渡すかを学ぶ切り替え方式である。段階選択器が現在の制御器を選び、安定化器が一時的な切り替えを抑え、行動境界の管理器がまとまった方策出力の間の移行を扱う。LIBEROの複数物体の移動・配置課題と、精密段階が複数ある実機のケーブル把持・ポート挿入課題で評価した。シミュレーションでは、学習した切り替えが基礎VLAを上回り、実行時に段階境界を知らなくても、その情報を利用できる切り替えと同等だった。実機評価では、操作者に合わせた引き継ぎ手順の下で、把持用と挿入用の両専門方策への自動切り替えを確認した。精密な適応が特に役立つところにRL専門制御を適用し、失敗した実行からの回復も含む汎用VLAの振る舞いを維持できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-language-action (VLA) models provide broad manipulation competence, but often struggle during the precision-critical stages that dominate contact-rich industrial tasks such as connector insertion and cable management. A common remedy is to refine a pretrained VLA with reinforcement learning (RL), enabling task-specific improvement beyond behavior cloning. However, how to preserve its generalist behavior while deciding when RL refinement is needed and which specialized policy should act remains an open question. In this work, we present RouteRLT, a routing framework that learns when and which RL specialist, an RL policy trained for a single precision-critical phase, should take control from a generalist VLA. A phase selector identifies the active controller, a stabilizer suppresses transient switches, and an action-boundary manager handles transitions between chunked policy outputs. We evaluate RouteRLT on multi-object pick-and-place tasks in LIBERO, as well as on a real-world cable pickup and port-insertion task with multiple precision-critical stages. In simulation, the learned routing improves over the base VLA and matches routing with privileged phase boundaries, without accessing those boundaries at deployment. The real-robot evaluation validates automatic routing to both the pickup and insertion specialists under an operator-aligned handoff protocol. Altogether, these results show that learned routing applies RL specialist control where precise adaptation is most valuable while preserving generalist VLA behavior, including recovery from failed execution attempts.

著者のコメント

8 pages, 7 figures, 2 tables. Accepted at the IROS 2026 International Workshop on Industrial Applications of Robot Learning (IARL)

arXiv ID: 2609.26467 / 要約の誤りについて