arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

言語モデルをロボットの操作方策として使う初期評価

An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond

Wenbo Zhang, Kaixuan Wang, Yutao Ouyang, Xiaoyu Huang, Liyang Li, Kailun Su, Weiyang Jin, Wenhao Chai, Haotian Liang, Zhiyang Dou, Yue Chen, Tianxing Chen

この論文をやさしく読む

ひとことで言うと

言語モデルに専用の操作学習を加えず、ロボットの行動を直接決めさせたときの成績を比較しています。

何に役立つ?

言語モデルを操作系に組み込む際、意味理解を生かせる作業と、精密な制御が不足する作業を見分ける材料になります。

この研究の面白いところ

公開方策を上回るScoreと、平均成功率22.48%という絶対的な成功水準を両方報告しています。単発の見本提示にも全体的な効果はありませんでした。

どこまで分かった?

モデル間で評価エピソード数が異なります。RoboDojo上の初期評価であり、一般の実機運用での信頼性は示していません。精密・動的・両手協調の制御に弱点があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

身体性を持つAIシステムは、システム1とシステム2に分けて構成されることが多い。システム1は通常、高頻度で動作を生成する事前学習済み方策であり、システム2は高水準の計画を行う画像対応言語モデルとして実装されることが多い。本研究では、タスク固有の追加学習なしに、大規模言語モデル(LLM)がロボット操作の方策として働けるかを問う。この設定を「方策としてのLLM」と呼ぶ。 三つのLLMをRoboDojoの全42タスクで評価し、公開されている40の方策と得点を比較する。AstraとGPT-5.5には、各タスク50エピソードの公式手順を用い、DeepSeek-Flashには各タスク10エピソードを用いる。GPT-6 Astraは2,100試行で平均成功率22.48%、Score 28.97を達成し、公開されているすべてのエントリーを上回る順位だった。一方、同じ後処理を用いても、GPT-5.5とDeepSeek-Flashの平均成功率は、それぞれ0.88%と1.92%にとどまった。 Astraの能力には強い二極化が見られる。意味理解を必要とするが高精度制御を必要としないタスクにはよく汎化する。対照的に、精密さ、動的制御、複雑な両手の協調を必要とするタスクでは性能が低い。文脈内学習の実験では、1件のデモンストレーションを与えても全体としての改善はなく、一方で選んだ対話記録には、摂動を受けた際に同一エピソード内で修正する様子が見られる。総じて、評価したLLMの操作性能には大きな差がある。Astraは目立った性能を示し、汎用的な操作モデルの可能性を示す初期的な証拠を提供するが、評価した設定では、信頼できる精密制御と動的制御が引き続き課題である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.

著者のコメント

24 pages

arXiv ID: 2609.24170 / 要約の誤りについて