arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

視覚言語モデルが動作を予行演習するロボット操作システム

World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal

Yehang Zhang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, Zexi Li

この論文をやさしく読む

ひとことで言うと

ロボットが動く前に、見えている場面で動作を試して修正する仕組みを、視覚言語モデルに与えた研究です。

何に役立つ?

考えられる用途は、ロボットの操作計画を実行前に見直すことや、操作記録を小型モデルの訓練に使うことです。要旨では複数のロボット操作ベンチマークで成功率を評価しています。

この研究の面白いところ

視点選択、動作の予行演習、実行時の位置ずれ補正を一つの作業空間にまとめています。LIBERO-90から得た技能が、追加学習なしでrobosuiteでも役立ったと報告しています。

どこまで分かった?

示された数値はLIBERO-Proなどの評価環境における結果です。実環境の多様なロボットや長期運用で同じ成功率になるかは、要旨からは分かりません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

汎用の視覚言語モデル(VLM)はロボット操作に幅広い知識と空間推論能力をもたらす。しかし既存のシステムは、制約の予測やプログラム作成にVLMを間接的に使うか、行動できる世界ではなく場面の画像をVLMに見せるにとどまっていた。本研究では、VLMが基本的なツールを使ってロボットを操縦し、すべての判断を視覚的な動作作業空間内で行う複数エージェントの枠組み、World Action Agent(WAA)を提示する。 この作業空間には三つの特徴がある。接触に関する視点は場面の幾何形状から自動選択され、現在の相互作用の周囲を示す。動作の予行演習では、各動作を編集可能な提案にし、エージェントが単独またはImagination Agentを通じて、実行前に計画上のフィードバックを基にプレビューと修正を行う。視野内での補正は観察、予行演習、低レベルの実行を結び、観察している視点内で残った位置ずれを除去できる。WAAは同じ作業空間を通じて、専門家の動画と人間の指導から証拠に基づく審査の下でマルチモーダルな技能を発展させ、Skill Agentを介して参照する。また、相互作用の記録を使って、より小さなVLMに同じ枠組みの操作を学習させる。 LIBERO-Proでは、LIBERO-90だけから発展させた技能を使うWAAの平均成功率は75.6%で、エンドツーエンドの視覚・言語・動作モデル、プログラムを方策として使うエージェント、同じ基盤モデルを使う視覚的枠組みの比較対象を上回り、最高水準の結果となった。同じ技能は追加学習なしでrobosuiteでも有効だった。枠組みの操作記録でQwen3.5-9Bを追加学習すると、分布外での成功率は1.7%から43.3%に上がった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.

著者のコメント

Working in progress

arXiv ID: 2609.29964 / 要約の誤りについて