arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

自由空間の移動を古典的計画に任せ、ロボット操作を高速化

SkipVLA: Skipping VLA Steps with Classical Planning for Fast Robot Manipulation

Kaivalya Agrawal, Md Ashiqur Rahman, Raymond A. Yeh, and Zachary Kingston

この論文をやさしく読む

ひとことで言うと

何も触らずに腕を移動させる部分は高速な動作計画器に任せ、つかむ・置く場面でVLAを使う方法です。

何に役立つ?

ロボットの操作時間と消費エネルギーの削減に役立ちます。シミュレーションだけでなく、実機アームの物体移動課題でも評価されています。

この研究の面白いところ

VLAが持つ視覚と言語の理解を目標姿勢の予測に使い、毎回の行動生成を省きます。追加の実演を導入せずに予測器を学習する点も特徴です。

どこまで分かった?

評価は三つのVLA、LIBEROの13課題、実機の三つの課題です。最大2.5倍は評価内の最大値で、すべての操作や環境で同じ短縮率が得られるという結果ではありません。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚・言語・行動(VLA)モデルは、カメラ画像と言語指示をロボットの行動へ直接対応付ける汎用的なロボット方策の一種である。有望ではあるが、実行時には依然として遅く、特に方策への問い合わせを何度も必要とする長期的な課題で問題になる。最近の研究では、小さいモデルへの蒸留、非同期の行動チャンクの重ね合わせ、高速な低水準方策との組み合わせによってVLAの遅延を削減しているが、それでも課題全体を通じて学習済みの方策を実行する。VLAとは対照的に、古典的な動作計画器は衝突のない動作を素早く見つけられるが、明示的な目標を必要とし、課題の意味を理解しない。 本研究では、事前学習済みVLAと古典的な動作計画器を組み合わせた混合方策SkipVLAを提示する。自由空間での動作には計画器を使い、把持や配置など、接触を多く伴う技能だけでVLAに問い合わせる。SkipVLAはVLAの凍結した視覚言語バックボーンを再利用し、計画する各動作の目標姿勢を予測する。大規模VLAがすでに学んだ内容を利用することで、システムに新たな実演データを追加せずに、この予測器を学習する。 三つのVLAを用い、シミュレーション内のLIBEROの13課題と、実機の6自由度YAMアームによる三つのピック・アンド・プレース課題で評価する。その結果、課題成功率を維持しながら、課題完了を最大2.5倍高速化し、エネルギー消費も大幅に低減した。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-18 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-Language-Action (VLA) models are a class of generalist robot policies that map camera images and language instructions directly to robot actions. While promising, these models remain slow at test time, particularly for long-horizon tasks that require many queries to the policy. Recent efforts reduce VLA latency by distilling smaller models, overlapping asynchronous action chunks, or pairing the VLA with a fast low-level policy, but still run a learned policy for the entire task. In contrast to VLA, classical motion planners quickly find collision-free motions, but require an explicit goal and have no semantic understanding of the task. In this work, we present SkipVLA, a hybrid policy that combines a pretrained VLA with a classical motion planner, using the planner for free-space motion and querying the VLA only for contact-rich skills such as grasping and placing. SkipVLA reuses the frozen vision-language backbone of the VLA to predict a target pose for each planned motion, and learns this predictor without additional demonstrations introduced into the system by using what was already learnt by the large VLA. We evaluate SkipVLA with three VLAs on 13 LIBERO tasks in simulation and three pick-and-place tasks on a physical 6-DoF YAM arm, demonstrating up to 2.5x faster task completion and significantly lower energy consumption while achieving the same task success rate.

arXiv ID: 2609.20648 / 要約の誤りについて