arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

ロボットの成功率だけでは見えない動作の変化を測る

Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks

Sophie Higham and Riccardo Andrea Izzo and Matteo Matteucci and Alessandro Suglia

この論文をやさしく読む

ひとことで言うと

ロボットが作業を成功させても、動きがぎこちなくなったり効率が変わったりしていないかを評価する研究です。

何に役立つ?

同じ成功率のモデルを比較するとき、動作の滑らかさや効率まで含めて選ぶための評価方法になります。

この研究の面白いところ

失敗だけを見るのではなく、成功した試行の中で起きる行動の変化とばらつきに注目しています。

どこまで分かった?

検証は3モデル、4タスク群、7摂動条件に限られます。要旨は行動指標の必要性を示しますが、実機の安全性や接触力の保証を行ったとは述べていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚・言語・行動(VLA)モデルは、ロボット操作タスクのベンチマークで高い成功率を達成してきた。近年は摂動に対する頑健性の評価が重視されているが、その頑健性も依然として主にタスク成功率(TSR)で測られている。本研究は、摂動下で成功した軌道がどのように実行されるかを特徴付け、モデルの行動上の頑健性を測る、ベンチマークに依存しない評価枠組みを提案する。 広く使われるLIBEROとLIBERO-Plusを拡張して、この方法を実装する。最先端のVLAモデル3つ、LIBEROの4つのタスク群、7つの摂動条件にわたり、動作の滑らかさ、効率、グリッパーの振る舞いなどの指標を含め、典型的な成功時の行動とそのばらつきの両方の変化を評価する。摂動は成功軌道の行動を変化させ得ることが分かったが、この現象は必ずしもTSRだけから推測できない。LIBEROの各タスク群で、同じ摂動条件の下、最先端のVLAモデルが同程度のTSRを達成しても、成功軌道上の振る舞いが大きく異なる場合を特定する。 したがって、タスク性能をより確かに評価するには、適切な頑健性指標は、タスクが完了したかだけでなく、完了する間にロボットがどう振る舞うかも捉えるべきだと論じる。VLAモデルの頑健性評価では、ロボットによる成功したタスク実行の性質とばらつきを特徴付ける行動評価指標を、TSRに併用できる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.

arXiv ID: 2610.01351 / 要約の誤りについて