arXiv論文メモ
新着一覧
cs.RO / cs.CV / cs.LG · 査読状況未確認

世界モデルで未知形状の部品挿入に対応するロボット

Generalizable Robotic Insertion with World Models

Nicklas Hansen, Iretiayo Akinola, Yijie Guo, Jie Xu, Bingjie Tang, Hao Su, Xiaolong Wang, Abhishek Gupta, Dieter Fox, Yashraj Narang

この論文をやさしく読む

ひとことで言うと

複数の挿入作業で学んだ世界モデルを使い、形状を知らない新しい部品をロボットが挿入できるか調べた研究。

何に役立つ?

部品ごとに専用の方策を作る負担を減らすロボット組立システムの設計に役立つ可能性がある。

この研究の面白いところ

最大90課題で学習し、未見の物体での追加学習なしの成功率は56%。世界モデルなしの比較手法は7%だった。

どこまで分かった?

成功率は研究で評価した未見の物体と挿入課題に対する値である。工場での長期運用やあらゆる形状への対応は要旨で実証されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

多品種のロボット組立には多様な部品を扱える適応的なシステムが必要だが、現在の手法は通常、挿入作業ごとに専用の方策に依存する。高い成功率には達しうるものの、新しい課題への導入には手間と時間がかかる。本研究は、ロボット自身の位置や動きなどの情報と、手首に取り付けたカメラからの生の視覚観測を組み合わせる世界モデルを使い、さまざまな部品挿入に対応する枠組みを提示する。 幾何形状の異なる最大90種類の挿入課題で単一の世界モデルを学習した。形状が未知の未見の物体では、追加学習なしで56%の成功率を達成し、世界モデルを使わない比較手法の7%を上回った。学習データに含める物体を増やすほど性能が向上し、拡張性も示した。さらに、学習から除外していた物体について汎用モデルを追加学習すると、初めから学習する場合よりデータ効率が大きく改善し、場合によっては十分な学習後の性能も上回った。著者らの知る限り、未知の物体を完全にデータ駆動で組み立てられる初のシステムであり、拡張可能で汎用的なロボット組立への重要な一歩だと位置付ける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robotic assembly in high-mixture settings requires adaptable systems that can handle diverse parts, yet current approaches typically rely on policies specialized to each insertion task. Although this can reach high success rates, it makes the process of deploying systems for new problems tedious and time consuming. We present a framework for generalizable insertion using world models that combine robot proprioceptive information with raw visual observations captured by a wrist-mounted camera. Our model-based approach trains a single world model on up to 90 insertion tasks with geometrically diverse parts, achieving 56% zero-shot success on unseen objects with unknown geometry compared to just 7% with a model-free baseline. Importantly, performance improves as more objects are included in the training dataset, demonstrating strong scalability. Lastly, finetuning the generalist model on held-out objects significantly enhances data-efficiency compared to training from scratch and, in some cases, achieves better asymptotic performance. To our knowledge, this is the first system capable of assembling unseen objects in an entirely data-driven manner, and thus represents a significant step toward scalable, generalizable robotic assembly systems.

著者のコメント

IROS 2026

arXiv ID: 2609.28258 / 要約の誤りについて