作業理解と動作生成を統合するロボットモデル
MachEmbodied-U0: Unified Understanding and Generation Model for Embodied Intelligence
この論文をやさしく読む
ひとことで言うと
ロボットが作業を理解し、操作する場所や場面の変化を予測しながら動作を作る統合モデルを提案した。
何に役立つ?
ロボット操作で、言語的な作業理解と視覚的な将来予測を制御へ結びつけるモデルの比較材料になる。
この研究の面白いところ
約4,200時間の実演で事前学習し、LIBEROで99.0%、LIBERO-Plusで82.5%の平均成功率を報告した。下流で直接教えていない小作業や操作場所のゼロショット予測も示した。
どこまで分かった?
数値は指定したベンチマークでの結果であり、現実世界の操作も検証したと記載されるが、その作業数や成功率は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
汎用ロボット制御には、作業の意図を理解し、どこに働きかけるかを特定し、場面がどう変わるかを捉え、精密な動作を生成するモデルが必要である。視覚・言語・行動モデルは強い意味的な事前知識を持つが、場面の動きを明示的にモデル化しないことが多い。世界モデルと行動を結ぶモデルは視覚予測と制御を組み合わせるが、細かな操作に必要な作業関連の意味的・空間的構造が表に出るとは限らない。本論文は、Mixture-of-Transformers構造で理解の専門モジュールと生成の専門モジュールを結ぶ、統合型の身体性基盤モデルMachEmbodied-U0(ME-U0)を提案する。小作業の予測と、行動できる場所の特定が、フローマッチングによる視覚的な場面変化と行動の同時生成を導く。視覚的な変化には、将来のRGB画像、奥行き、表面法線、オプティカルフローを含め、見た目、幾何、運動について補完的な教師信号を与える。複数の時間間隔に対応する回転位置符号化(MRPE)が、視覚的変化と細かな制御を対応づける。 ロボットのデータセットと一人称視点のデータセットから選んだ、約4,200時間の実演でME-U0を事前学習した。各下流ベンチマークに元から備わる教師信号だけを用いると、RoboDojoのシミュレーションベンチマークで平均スコア17.66、LIBEROで平均成功率99.0%、LIBERO-Plusで82.5%を達成した。現実世界のロボット操作作業でも追加検証し、シミュレーション以外での有効性を示した。対応する下流の教師信号を与えなくても、シミュレーションと実世界の観測に対して、小作業の予測、行動可能な場所の特定、視覚的変化のゼロショット予測も示した。全体としてME-U0は、競争力のある下流制御性能と、シミュレーションおよび実世界の間で移せる作業の位置づけ・視覚的変化の能力を組み合わせる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual prediction with control without necessarily exposing the task-relevant semantic and spatial structure needed for fine-grained manipulation. We present MachEmbodied-U0 (ME-U0), a unified embodied foundation model connecting understanding and generation experts through a Mixture-of-Transformers architecture. Subtask prediction and affordance grounding guide joint visual-dynamics and action generation via flow matching. Visual dynamics encompass future RGB, depth, surface normals, and optical flow, providing complementary supervision for appearance, geometry, and motion. Multi-rate Rotary Position Encoding (MRPE) aligns visual dynamics with fine-grained control. We pretrain ME-U0 on approximately 4,200 hours of curated demonstrations from robotic datasets and egocentric datasets. Using only the supervision natively available in each downstream benchmark, ME-U0 achieves an average score of 17.66 on the RoboDojo simulation benchmark and average success rates of 99.0\% and 82.5\% on LIBERO and LIBERO-Plus, respectively. We additionally validate ME-U0 on real-world robotic manipulation tasks, demonstrating its effectiveness beyond simulation. Without corresponding downstream supervision, ME-U0 further demonstrates zero-shot subtask prediction, affordance grounding, and visual dynamics on simulated and real-world observations. Overall, ME-U0 combines competitive downstream control performance with transferable task-grounding and visual-dynamics capabilities across simulation and the real world.
著者のコメント
Technical report. Project page: https://machembodied.com/ME-U/ME-U0.html. Code: https://github.com/MachEmbodied/ME-U0
arXiv ID: 2609.25627 / 要約の誤りについて