arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

視覚言語モデルにロボットを逐次操作させる方法

Transferring the Intelligence of VLMs to Robotic Control

Meng-Hao Guo, Zhe-Han Mo, Jia-Jun Wang, Yi Zhang, Kejin Wang, Yi-Xuan Deng, Jia-Peng Zhang, Yongming Rao, Shi-Min Hu

この論文をやさしく読む

ひとことで言うと

視覚言語モデルに短い移動・回転・把持命令を渡し、観測と動作を繰り返してロボットを操作させた。

何に役立つ?

課題専用のロボット学習を減らし、少数の実演を使って新しい操作へ適応させる方法として参考になる。

この研究の面白いところ

一つの実演でRoboTwin 2.0 C2Rの成功率が53.2%から73.6%へ上がり、実機でも二種類の作業を行った。

どこまで分かった?

要旨での実機実証はFrankaによるブロック作業である。幅広い実世界作業での成功率は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人間は物理世界とデジタル世界の双方に柔軟に適応できる。このことは、身体、環境、課題には両世界の隔たりがあっても、知能そのものは移転できる可能性を示唆する。視覚言語モデル(VLM)の知能もデジタル世界から物理世界へ一般化し、ロボットを制御できるかを調べる。RoboDawnは、移動、回転、グリッパー操作を表す少数の離散的な命令を通して、エージェント型VLMにロボット制御を与える、人間に分かりやすい操作方法である。VLMは現在の視覚状態を観測し、次の動作を考えて実行し、その結果に応じて次の判断を変える閉ループでロボットを制御する。さらに、少数の実演を使い、操作方法と課題解決の戦略をVLMに身に付けさせる文脈内学習法を導入する。RoboTwin 2.0 C2RとRoboDojoでの実験では、課題専用のロボット学習なしで高い性能を示した。実演なしの場合も、ベンチマーク専用のロボットデータで学習した複数の強い方策を上回り、一つの実演を文脈内で与えるとさらに大きく改善して、最高水準の結果となった。RoboTwin 2.0 C2Rでは成功率が実演なしの53.2%から実演一つで73.6%へ上がり、比較手法π0.5の46.0%を超えた。RoboDojoでも35.67%から47.17%へ改善した。同じ枠組みは実機のFrankaロボットにも移り、ブロックをかごへ入れる作業とブロック積みを行った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Humans can seamlessly adapt to both physical and digital worlds, suggesting that while a digital-to-real gap exists in embodiment, environment and task, human intelligence itself may transfer across this gap. This naturally raises a fundamental question: can the intelligence of vision-language models (VLMs) similarly generalize from the digital world to the physical world for robotic control? We investigate this question through RoboDawn, a human-intuitive interface that exposes robotic control to an agentic VLM through a compact set of discrete translation, rotation, and gripper commands. Using this interface, the VLM controls a robot in a closed loop: it observes the current visual state, reasons about the next action, executes it, and adapts subsequent decisions to the resulting state. Furthermore, we introduce an in-context learning (ICL) scheme that uses a few demonstrations to ground the VLM in both interface usage and task-solving strategies. Experiments on RoboTwin 2.0 C2R and RoboDojo demonstrate that RoboDawn achieves strong performance without task-specific robot training. In the zero-shot setting, RoboDawn outperforms several strong policies trained on benchmarkspecific robot data, while a single in-context demonstration further yields substantial performance gains and establishes state-of-the-art (SOTA) results. On RoboTwin 2.0 C2R, the success rate increases from 53.2% zero-shot to 73.6% one-shot, exceeding the solid baseline {\pi}0.5 (46.0%). Similar gains are observed on RoboDojo, where success rate improves from 35.67% zero-shot to 47.17% one-shot. The same framework also transfers to real-world robots, performing block-in-basket and block stacking on Franka.

著者のコメント

For more detail, https://robodawn.top/

arXiv ID: 2609.22966 / 要約の誤りについて