arXiv論文メモ
新着一覧
cs.RO / cs.CV · 査読状況未確認

位置と姿勢の軌道で多様な動きを共通に学ぶWMM

World Motion Models: Flexible Sequence Modeling of SE(3) Trajectories

Jiahui Lei and Qianqian Wang and Trevor Darrell and Angjoo Kanazawa

この論文をやさしく読む

ひとことで言うと

物体や人体、ロボットの動きを位置と姿勢の軌道にそろえ、同じモデルで予測や補完、制御を扱う方法です。

何に役立つ?

動作ごとに別の表現を用意せず、異なる種類の動きやタスクを共通のモデルで学習するために役立つ可能性があります。

この研究の面白いところ

何を既知として与え、何を生成させるかをマスクで変えることで、多様なタスクを同じネットワークへの条件付けとして扱います。

どこまで分かった?

動的場面を剛体軌道の集合で近似する表現に基づきます。6用途での実験が報告されていますが、要旨には各用途の定量値や比較対象の詳細は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人工エージェントに空間的知能を持たせるには、動的な3次元世界について包括的な生成事前分布が必要である。本研究では、疎なSE(3)姿勢軌道を使い、何が過去・現在・未来のどの時点でどこにあるかを捉えるWorld Motion Models(WMM)を提案する。 WMMは、動的場面の要素が剛体のSE(3)軌道の集合でよく近似できるという観察に基づく。これは4次元のモデル化に必要最小限でありながら表現力のある基本要素である。この表現は、関節を持つ物体、人体、手と物体の相互作用、部分ごとに剛体とみなせる場面の動き、カメラの動き、さらにはロボットの状態と行動まで、一つの共通空間に統合する。 この表現の下で、これらの対象の同時分布を、トークンごとに雑音水準を持つフローマッチングを利用した柔軟な系列モデル化の問題として定式化する。系列ではない条件情報を与える文脈トークンの仕組みと組み合わせることで、任意個数の対象と時間ステップにまたがり、任意の一部を条件として別の任意の一部を扱う条件付けを支える。将来予測、動きの欠落補完、モデル予測制御、逆運動学、異なる身体構造への動作移し替え、方策学習などのタスクは、同じネットワークに異なるマスクを適用することへ帰着する。3次元視覚とロボティクスの多様な6用途での実験により、WMMの汎用性と柔軟性、および高い性能を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Equipping artificial agents with spatial intelligence requires a comprehensive generative prior over the dynamic 3D world. We propose World Motion Models (WMMs) that capture "what was, is, and will be where across time" via sparse SE(3) pose trajectories. WMMs are built on the observation that elements of dynamic scenes can be well approximated by a set of rigid SE(3) trajectories, a minimal yet expressive primitive for 4D modeling. This representation unifies articulated objects, human bodies, hand-object interactions, piecewise-rigid scene dynamics, camera motion, and even robot states and actions into a single shared space. Given this representation, we cast the joint distribution of these entities as a flexible sequence modeling problem, utilizing flow-matching with per-token noise levels. Coupled with a context token mechanism for non-sequential conditioning, this formulation supports any-to-any marginal conditioning across an arbitrary number of entities and time steps. Tasks such as future prediction, motion infilling, model-predictive control, inverse kinematics, cross-embodiment retargeting, and policy learning all reduce to the application of different masks over the same network. Experiments on 6 diverse applications of 3D vision and robotics demonstrate the versatility and flexibility of WMMs with strong performance.

著者のコメント

Accepted at NeurIPS 2026 (Spotlight). Url: https://jiahuilei.com/projects/wmm/

arXiv ID: 2610.01742 / 要約の誤りについて