arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

奥行きや把持候補を道具で得るロボット操作方式

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Zexi Li, Yehang Zhang, Wenqian Li, Haojian Huang, Chenxu Wang, Shiyuan Deng, Yangkai Wei, Tianyi Zhang, Binghui Xie, Bohan Zhou, Yifan Chang, Kaiwen Zhou, Ying-Cong Chen, James Cheng, Yinchuan Li

この論文をやさしく読む

ひとことで言うと

視覚言語モデルに奥行き、位置、把持候補を調べる道具を渡し、その情報からロボットの動作を選ばせる方式です。

何に役立つ?

ロボット操作で、モデル自体を変更せず三次元情報を使うための設計例になります。報告された効果はLIBERO-PRO、RoboSuite、RoboTwinでの評価に基づきます。

この研究の面白いところ

同じK1を用いるとAstraの正解率が88.9%になり、107エピソードから学んだ生徒モデルも比較対象のOpenVLAを上回りました。

どこまで分かった?

RoboTwinの成功率はEasyで32.0%、Hardで28.0%であり、すべての操作に成功したわけではありません。現実環境全般への性能は要旨からは判断できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

基盤的な視覚言語モデル(VLM)は物体、指示、空間関係を理解できるが、その能力をロボットの物体操作に結び付けるのは難しい。視覚言語行動モデル(VLA)は大量の実演を必要とし、事前学習で得た理解を損なう可能性がある。一方、RGB画像だけを使ってVLMが直接制御する方式はコストが高く、モデルの能力に強く依存する。本研究は、知覚をツールとして公開するロボット利用エージェントの枠組みRobo-Harness K1を導入する。エージェントは、較正済みの奥行き情報、継続的に追跡する視覚上の目印、空間的な測定値、把持候補を問い合わせ、その結果に基づいて汎用的な動作を選ぶ。このインターフェースによって、VLMの構造を変更したり奥行きエンコーダを学習したりせずに、三次元の幾何情報を利用できる。 条件を揃えたLIBERO-PROタスクでは、K1を使うGemini 3.7 Flashの正解率は77.8%となり、RGBのみの実行枠組みを使うGPT-6 Astraの61.1%を上回った。K1を使うとAstraも88.9%に改善した。対象タスクでの追加微調整なしに、GeminiとK1の組み合わせは3種類のRoboSuiteロボットアームと、双腕のRoboTwinタスクへ転移した。RoboTwinではEasyで32.0%、Hardで28.0%を達成し、視覚や環境の変動への耐性を示した。K1は次トークン学習に適したツール呼び出しの記録も生成する。教師役の107エピソードだけで学習したQwen3.5-9Bの生徒モデルは、新しい初期状態で44.2%の正解率となり、OpenVLAの30.2%を上回った。学習時に使わなかったタスク条件でも13.9%で、OpenVLAの0.0%を上回った。これらの結果は、知覚ツールを加えたロボット利用エージェントが、VLMの能力を使って、少数の例から学べる汎化可能なロボット方策につながる可能性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Foundation vision-language models (VLMs) understand objects, instructions, and spatial relations, yet translating this capability into robotic manipulation remains difficult. Vision-language-action (VLA) models require extensive demonstrations and may compromise pretrained understanding, while direct RGB-only VLM control is costly and strongly dependent on model capability. We introduce Robo-Harness K1, a robot-use agent (RUA) framework that exposes perception as tools. The agent queries calibrated depth, persistent visual anchors, spatial measurements, and grasp hypotheses, then selects generic motions from the returned evidence. This interface makes 3D geometry accessible without changing the VLM architecture or training a depth encoder. On matched LIBERO-PRO tasks, Gemini 3.7 Flash with K1 reaches 77.8% accuracy, surpassing GPT-6 Astra's 61.1% with an RGB-only harness; K1 further improves Astra to 88.9%. Without target fine-tuning, Gemini with K1 transfers to three RoboSuite arms and dual-arm RoboTwin tasks. On RoboTwin, it achieves 32.0% on Easy and 28.0% on Hard, showing resilience to visual and environmental perturbations. K1 also produces tool-call traces aligned with next-token training. A Qwen3.5-9B student trained on only 107 teacher episodes reaches 44.2% accuracy on new initial states versus 30.2% for OpenVLA, and 13.9% on held-out task conditions versus 0.0% for OpenVLA. These results suggest that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies that leverage VLM capabilities through an accessible tool interface.

著者のコメント

preprint

arXiv ID: 2609.29389 / 要約の誤りについて