arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

公開ロボット方策6種を共通条件で比較する

IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models

Yiqi Wang, Zhifeng Rao, Jiaqi Zhang, Xiaoyang Li, Zhangkai Wu, Yiqun Duan, Mingkai Zheng, Fei Wang, Shan You, Taotao Cai

この論文をやさしく読む

ひとことで言うと

公開ロボットモデル6種を、操作能力、条件変化への強さ、指示文への感度、計算コストに分けて比べた。

何に役立つ?

ロボット方策の採用や研究上の比較で、通常条件の成功率だけでなく頑健性と実行コストを確認するための評価基準になる。

この研究の面白いところ

通常条件の平均差は小さい一方、頑健性や言い換えへの対応には大きな差があった。比較の根拠の強さを区分し、厳密な比較対象を限定している。

どこまで分かった?

6つの公開システムと示されたLIBERO系の評価条件に基づく。要旨はVLAとWAMのどちらかが常に優れるとは述べていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

公開されているロボットの方策は、主に2つの方式に分かれる。視覚・言語・行動モデル(VLA)は観測と指示を直接行動へ写像し、世界・行動モデル(WAM)は学習した動画や世界の動きを方策学習または行動生成に取り入れる。両者は同じ物体操作課題を対象とする代替設計だが、通常は異なる評価手順で報告されるため、能力、頑健性、言語表現への感度、実行コストの得失が分かりにくい。 著者らはIndustrialVLA-Benchを提示し、公開されたVLAとWAMの6システムを、根拠の状態が分かる共通の報告形式で評価する。LIBERO上の通常条件での能力、LIBERO-Plus上の言語以外の変化への頑健性、LIBERO-Para上の指示文への感度、観測された実行コストを別々に評価した。課題スコアは、チェックポイントと推論設定を固定し、異なる乱数種を使った3回の完全な評価を集約した。 6システム全体では、通常条件のLIBERO平均値の幅はわずか1.58ポイントだったのに対し、頑健性と指示文の言い換えに関する集計の幅はそれぞれ14.62、31.08ポイントだった。比較対象を、手順への忠実性が確認できた3システムに限定しても、それぞれ1.36、14.62、23.10ポイントという差が保たれた。このため、報告された評価軸間の違いは、根拠の弱い区分に依存しない。各システムについて、観測された推論遅延、最大メモリ使用量、実行モード、根拠の状態も報告する。手順に忠実な実施、ほぼ再現した実施、検証待ちを明確に分け、厳密な比較は手順に忠実な実施に限る。どちらかの方式が普遍的に優れると主張するのではなく、共通の実用的基準で公開ロボット方策を比較できる、追跡可能な根拠を提供する。コードと評価記録は論文に記載のGitHubで公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate learned video or world dynamics into policy learning or action generation. Although both target the same manipulation tasks and represent alternative design choices, they are commonly reported under different evaluation protocols, leaving their capability, robustness, language sensitivity, and deployment-cost trade-offs unclear. We present IndustrialVLA-Bench, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema. It separately evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost. Reported task scores aggregate three complete evaluations with distinct random seeds under a fixed checkpoint and inference configuration. Across all six systems, clean LIBERO averages differ by only 1.58 points, whereas robustness and paraphrase summaries span 14.62 and 31.08 points. Restricting every comparison to the three protocol-faithful systems preserves the effect (1.36, 14.62 and 23.10 points), so the diagnostic separation reported here does not depend on the weaker evidence tiers. We additionally report observed inference latency, peak memory, runtime mode, and an evidence status for every system. Protocol-faithful, near-reproduction, and pending-verification entries remain visibly separated; only protocol-faithful entries support strict comparisons. Rather than claiming universal superiority of either paradigm, IndustrialVLA-Bench provides traceable evidence for comparing released robot policies on shared practical criteria. Code and evaluation records are available at https://github.com/xiaoqi-7/IndustrialVLA-Bench.

著者のコメント

preprint

arXiv ID: 2609.25562 / 要約の誤りについて