arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

変化する現実環境で動画予測とロボット制御を両立するInternW0

InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

Jisong Cai, Yao Mu, Ganlin Yang, Zhe Cao, Zhangzheng Tu, Xing Gao, Kailin Li, Xinyu Zhan, Lixin Yang, Yangkun Zhu, Haoxiang Ma, Ming Zhou, Qiaojun Yu, Yufei Xue, Liqun He, Yifei Yao, Yifan Zhu, Long Ling, Bingqi Jiang, Haoyu Guo, Xueyue Zhu, Bowen Zhou, Bin Zhao, Tianfan Xue, Chunhua Shen, Weinan Zhang

この論文をやさしく読む

ひとことで言うと

動画による先読みと素早いロボット操作を組み合わせ、実験室の作業まで扱う世界モデルを提案した。

何に役立つ?

考えられる用途は、観測が変わり続ける環境でのロボット操作や科学実験の自動化である。要旨にある評価はシミュレーションと特定の実作業で行われた。

この研究の面白いところ

長期予測を担当する大きなモデルと高速制御を担当する小さなモデルを分け、予測情報を再利用する構成が特徴である。

どこまで分かった?

約7,200時間のデータで学習し、15段階の合成作業などで評価した。普遍的な実世界作業への適用可能性を全て実証したとは述べていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

物理的な知能には、世界がどう変化するかを予測するだけでなく、世界が変化し続ける中でもその予測を行動に使えることが必要である。上海AI研究所の物理世界モデル群InternWの最初の実装としてInternW0を提案する。多様な入出力形式、非同期で頻度の異なる処理、部分的な観測や外部からの影響のもとでの局所的な物理モデリングを中心に設計した。 InternW0は、フローマッチングを使う非対称な動画・行動アーキテクチャにより、将来の視覚的変化と連続的なロボット制御を共同で学習する。大容量の動画専門モデルがより長い時間幅の予測情報を提供し、軽量な行動専門モデルはより速い時間尺度で動作する。行動を更新するたびに将来を生成し直す代わりに、各層のキー・バリュー情報を再利用し、新たに観測した状態に合わせて観測条件付きの文脈ルーティングで適応させる。領域別のインターフェースとソフトプロンプトが異なるロボット形態に対応し、接触を考慮した追加学習には力と触覚の信号を取り入れる。 EgoLabという実験室での一人称視点データ275時間分を含む、異種のロボット・一人称視点データ約7,200時間分で学習した。評価にはシミュレーションのベンチマークと実世界の科学作業を含め、有機金属構造体の合成における15段階の手順と、汎用的な定量ピペッティングに向けた接触・力を考慮する5段階の器用な操作を扱った。これらの結果は、拡張可能で非同期に動作し、科学作業を扱う物理世界モデルによる、効率的な実世界との相互作用に向けた進展を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences. InternW0 jointly learns future visual dynamics and continuous robot control through an asymmetric video--action architecture with flow matching. A high-capacity video expert provides longer-horizon predictive context, while a lightweight action expert operates at a faster timescale. Instead of regenerating the future for every action update, InternW0 reuses layerwise K/V and adapts it to newly observed states through observation-conditioned context routing. Domain-specific interfaces and soft prompts support heterogeneous embodiments, while contact-aware post-training incorporates force and tactile signals for contact-rich manipulation. We train InternW0 on approximately 7,200 hours of heterogeneous robot and egocentric data, including EgoLab, a 275-hour real-laboratory egocentric dataset. Evaluation spans simulation benchmarks and real-world scientific tasks, including a 15-stage metal--organic framework synthesis workflow and 5-stage contact- and force-aware dexterous manipulation for general-purpose quantitative pipetting. These results advance scalable, asynchronous, and science-native physical world models for universal and efficient real-world interactions.

著者のコメント

A technical report of world models, 24 pages, 8 figures, and 7 tables

arXiv ID: 2609.27656 / 要約の誤りについて