ロボットの指示理解を測るRoboFollow
RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents
この論文をやさしく読む
ひとことで言うと
ロボットが高得点を取っていても、言葉を理解して動いたとは限らないため、指示の理解を切り分けて測るベンチマークを作った。
何に役立つ?
ロボットの視覚・言語・行動モデルが、場面から作業を推測しているだけか、言語指示を使っているかを評価するのに役立つ。
この研究の面白いところ
一つの場面に複数の作業を設定して言語を必要にし、意図の理解と動作の実行も別の得点で示す。9種類の方策では、基本設定での好成績が難しい設定へ安定して移らなかった。
どこまで分かった?
結果は用いた微調整設定、9種類の方策、L0~L3の診断条件に基づく。要旨に挙げた対策では差を埋められなかったが、あらゆる学習法について調べたわけではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
近年の身体性を持つエージェントは高い成功率を達成しているが、実際に指示に従う能力は、その数値が示すほど強くない。本研究は、この見かけ上の成功の原因を、視覚的な場面に有効な作業が一つしかない「場面のエントロピーが低い」状態に求める。この場合、言語情報は余分になり、方策は言語をほとんど使わずに高得点を取れる。診断用ベンチマークRoboFollowは、三つの原則に基づく。第一に、各訓練場面で運動学的に異なる複数の作業分岐を用意し、視覚だけでは不十分にして、言語への依存を必要にする。第二に、4段階の診断手順(L0~L3)で視覚配置と意味を段階的に変え、空間関係、属性、軌道の制約、論理にわたって、同等の指示には一貫した動作を、異なる指示には区別できる動作を示すか調べる。第三に、交絡を抑えるため、操作対象を単純化し、動作を訓練済みの範囲に限り、意図と実行の得点を段階別に報告して、理解と運動実行を切り分ける。 9種類の視覚・言語・行動(VLA)方策とWAM方策の評価では、L0で高い成績に達しても、用いた微調整設定の下ではL1~L3へ安定して移らなかった。より強い視覚言語モデルを基盤にする方法、質問応答を併用した学習、LangForce、Classifier-Free Guidanceといった代表的な対策でも、この差は埋まらなかった。RoboFollowは、実際に指示に従うことが重要で、見落とされがちな障壁だと示す。コードとデータセットが公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.
arXiv ID: 2609.25636 / 要約の誤りについて