arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

人の一人称映像をロボットの事前学習に生かす条件を探る

AtomEgo: Exploring Ego-Robot Integration for Embodied Foundation Model Pretraining

Di Wu, Dongchen Zheng, Junhe Sheng, Zhongxing Wei, Songxin Zhang, Zejian Xie, Xiaoquan Sun, Junyang Zheng, Zhuoyang Song, Jiaxing Zhang, Jiayu Chen

この論文をやさしく読む

ひとことで言うと

人が物を扱う一人称映像を、ロボットの学習にどう組み込むと役立つかを三つの方式で比較する研究です。

何に役立つ?

ロボットの実演だけではデータが足りないときに、人のデータを使う設計の参考になります。データを増やすことに加え、人とロボットの身体・行動の違いを整えることが重視されています。

この研究の面白いところ

約2,659時間のデータを使い、異なるモデル構造にまたがって共同学習と段階的転移を調べます。実ロボットの複数タスクと表現分析を組み合わせています。

どこまで分かった?

要旨にはタスク別の改善量や詳細な実験条件はありません。「データ規模 × 整合の質」は結果をまとめた原則であり、能力向上を数値で予測する厳密な法則として提示されているわけではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

身体を持って行動する基盤モデルは、ロボットの実演データの規模と多様性が限られていることに制約されており、大規模な人間の一人称視点のインタラクションデータを使う動機となっている。しかし、人間とロボットの身体および行動空間には大きな隔たりがあるため、こうしたデータを身体性を持つモデルの事前学習にどう効果的に取り入れるかは明確でない。 本研究では、精選した約2,659時間のコーパスと拡張可能なデータ処理パイプラインに支えられた、人間の一人称データとロボットデータの共同学習に関する体系的な研究、AtomEgoを提示する。視覚・言語・行動モデルと世界・行動モデルのアーキテクチャにわたり、代表的な三つの方式を調べる。すなわち、ドメイン別の行動ヘッドによる共同学習、身体の違いを整合させながら人間の一人称データからロボットへ段階的に転移する方式、動画と行動の共同モデリングである。 複数タスクの実ロボット実験と、言語を条件とする異なる身体間の表現分析を通じて、これらの方式を評価する。結果は「データ規模 × 整合の質 → 能力向上」という単純な原則を明らかにする。一人称データは汎化を改善できるが、その価値はデータをどれほど効果的に整合させ、活用できるかに依存する。この原則は、拡張可能な人間一人称・ロボット共同事前学習に実践的な指針を与え得る。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Embodied foundation models are constrained by the limited scale and diversity of robot demonstrations, motivating the use of large-scale egocentric human interaction data. However, how to effectively incorporate such data into embodied-model pre-training remains unclear because of substantial embodiment and action-space gaps between humans and robots. We present AtomEgo, a systematic study of ego--robot co-training supported by a curated corpus of approximately 2,659 hours and a scalable data processing pipeline. Across vision--language--action and world--action model architectures, we investigate three representative paradigms: joint co-training with domain-specific action heads, progressive ego-to-robot transfer through embodiment alignment, and joint video--action modeling. We evaluate these paradigms through multi-task real-robot experiments and language-conditioned cross-embodiment representation analysis. Our results reveal a simple principle: Data Scale * Alignment Quality --> Capability Gain; egocentric data can improve generalization, but their value depends on how effectively they are aligned and utilized. This principle can provide practical guidance for scalable ego--robot pre-training.

arXiv ID: 2609.21461 / 要約の誤りについて