3次元潜在粒子による物体中心のシーン表現
ParticleSplat: Self-supervised Object-centric Latent Particle Splatting
この論文をやさしく読む
ひとことで言うと
複数方向からの画像を、物体に対応する3次元の潜在粒子として表現する方法です。粒子を動かすことで、物体を動かした場面も構成できます。
何に役立つ?
ロボットが物体の位置や形を扱うための表現として役立ちます。要旨では、学習した表現によるロボット操作タスクの性能改善を報告しています。
この研究の面白いところ
従来の2次元の粒子表現と3次元ガウス表現の構造的な近さを利用しています。新しい視点の画像を再構成する学習から、物体マスクも教師なしで得ています。
どこまで分かった?
シミュレーションと実世界のデータセットで評価しています。入力には複数視点とカメラ姿勢を用い、要旨には操作性能の改善幅や個々の失敗条件は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ParticleSplatは、シーンを意味的な実体を表す一組の潜在的な「粒子」に分解する、自己教師ありの物体中心学習手法です。画像を位置、スケール、見た目などの属性をもつ粒子の集合として表すDeep Latent Particles(DLP)に基づきつつ、DLPが本質的に2次元であるため、ロボット操作などの下流タスクで重要な3次元の空間・幾何推論を明示的に行えないという制限に対処します。 潜在粒子と3次元ガウスプリミティブの構造的な類似性を利用し、新しい視点合成目的関数で学習する3次元潜在粒子空間を導入します。モデルはカメラ姿勢とともに複数視点を共通の3次元物体中心潜在空間へ符号化し、その後、粒子を粒子に対応づけられた3次元ガウスへ変換します。これらを合成することでシーン全体を再構成します。 シミュレーションデータセットと実世界データセットで評価した結果、この定式化は教師なしで物体マスクを本質的に学習し、潜在空間の粒子を変更することで物体を移動させるなど、制御可能な3次元シーン編集を可能にしました。さらに、学習された3次元表現がロボット操作タスクの下流性能を改善することを示しました。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which represents images as a set of particles with attributes such as position, scale, and visual appearance, we address a key limitation of DLP: its inherently 2D nature, which prevents explicit 3D spatial and geometric reasoning that are critical for downstream tasks such as robotic manipulation. Leveraging the structural similarity between latent particles and 3D Gaussian primitives, we introduce a 3D latent particle space trained with a novel view synthesis objective. Our model jointly encodes multiple views with camera poses into a shared 3D object-centric latent space, then transforms particles into particle-aligned 3D Gaussians whose composition reconstructs the full scene. On simulated and real-world datasets, we show that this formulation inherently learns object masks without supervision and supports controllable 3D scene editing, such as moving objects by modifying particles in the latent space. We further establish that the learned 3D representation improves downstream performance on robotic manipulation tasks.
著者のコメント
Project page: https://lyuxinghe.github.io/ParticleSplat-website/
arXiv ID: 2609.19463 / 要約の誤りについて