硬い物体・布・ロープを扱うロボット用の点群世界モデル
PointCast: One World Model for Rigid, Articulated, and Deformable Object Manipulation
この論文をやさしく読む
ひとことで言うと
物体上の3次元点の軌跡を予測し、剛体から布やロープまで扱うロボット用世界モデルを提案した。
何に役立つ?
考えられる用途は、多様な物体の操作を実行前に予測・計画すること。要旨にはシミュレーションと実ロボットの遠隔操作データでの評価がある。
この研究の面白いところ
メッシュや接続構造に依存せず、各点の識別と軌跡を保ちながら学習する。
どこまで分かった?
シミュレーションの4種類中3種類で最良、実データ6分類中4分類で最小誤差だった。計画性能の評価は要旨ではシミュレーション上の64エピソードである。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
世界モデルは、ロボットが行動を実行する前に、物体の状態がどう変わるかを予測するために役立つ。本研究は、剛体、関節のある物体、変形する物体の操作にまたがって使える点集合の世界モデルPointCastを提示する。状態は物体とエンドエフェクタ上の3次元点の集合であり、メッシュを必要とせず、物体の接続構造にも依存しない。各点は自身の識別子を保ち、その軌跡に対して個別に教師信号を受ける。このため、点群が全体として作る形だけでなく、各点がどこへ動くかを学ぶ。 基盤には拡散トランスフォーマーを用い、点の最近の履歴と指示されたエンドエフェクタの動きを条件として、将来の短い期間の点位置のノイズを除く。注意機構は局所と大域を交互に使い、エンドエフェクタへの交差注意が両者の結び付きを表す。1,980万パラメータの一つの構成と一つの学習手順で、剛体、布、ロープ、多関節キャビネットの4種類を扱い、それぞれに別のチェックポイントを学習した。 ランダム化したシミュレーションで学習し、同じ指標で4つの比較対象と比べると、4種類中3種類で最良、剛体では2位だった。実ロボットの遠隔操作データで学習した場合、6分類のうち4分類で平均誤差が最小、残る2分類では2位となり、6分類すべてでそのデータセットに付属するモデルを改善した。シミュレーションで学習したチェックポイントを追加学習なしで適用した場合は、4つの収録例中2つで最良だった。モデルを固定し、期間当たり1回のネットワーク評価でサンプリング型モデル予測制御に組み込むと、シミュレーション上の4タスク、64エピソードで計画を行い、各タスクで全比較対象と同等以上の結果だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
World models are useful for robotic manipulation because robots can predict how actions change the states of objects before executing them. We present PointCast, a point-set world model that spans rigid, articulated, and deformable object manipulation. Its state is a set of 3D points on the object and the end-effector, mesh-free and topology-agnostic. Each point keeps its identity and is supervised on its own trajectory, which teaches the model where every point goes rather than only the shape the points form. Its backbone is a diffusion transformer that denoises a short window of future point positions, conditioned on the points' recent history and the commanded end-effector motion. The backbone's attention alternates between local and global, and cross-attention to the end-effector carries the coupling. This one architecture at 19.8M parameters and one training recipe cover four regimes, rigid objects, cloth, rope, and multi-joint cabinets, with a separate checkpoint trained for each. Trained on randomized simulation and scored against four baselines on the same metric, it is best on three of four regimes and second on rigid. Trained on a real-world robot teleoperation dataset, it has the lowest mean error in four of its six categories, is second in the other two, and improves on the dataset's own model in all six; zero-shot, its simulation checkpoints are best on two of four captures. Frozen inside sampling-based model-predictive control at one network evaluation per window, it plans four simulated tasks over 64 episodes, competitive with or outperforming every baseline on each. Project website at https://pointcast-wm.github.io.
著者のコメント
8 pages, 8 figures, 5 tables. Project page: https://pointcast-wm.github.io. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
arXiv ID: 2609.28393 / 要約の誤りについて