動画1本の撮影からロボットの学習用実演を合成する
Demonstration Synthesis from a Single Scan via Gaussian Splatting for Visuomotor Policy Learning
この論文をやさしく読む
ひとことで言うと
作業場所を動画で一度撮り、その3次元の複製からロボットの練習例を大量に作る方法です。合成した実演だけで学習し、実機でも評価しています。
何に役立つ?
考えられる用途は、場所に合ったロボットの実演データを人が何度も収集する負担の削減です。実世界のUR10eで84.2%の成功率を報告しています。
この研究の面白いところ
画像の忠実度は3DGSで確保し、動作計画は運動学で作ります。接触が結果を左右する把持形成だけに学習済みの接触モデルを使います。
どこまで分かった?
95.1%は実世界の代わりに使ったシミュレーション場面、84.2%が実機の結果です。物理エンジンは生成ループにありませんが、事前学習された接触モデルの利用はあります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚運動方策の学習には、対象環境によく一致する大量の実演が必要だが、その都度新たに収集するのは依然として高価である。既存の実演合成法はこの費用を減らすものの、多大な手作業、視覚的な忠実度の限界、物理シミュレータへの強い依存に制約されている。本論文では、1回の動画スキャンだけを人間からの入力とし、生成ループに物理エンジンを使わずに実演を大量生成する、高忠実度のデータ生成基盤GaussianFactoryを導入する。 具体的には、場面を編集可能な3次元Gaussian Splatting(3DGS)の複製として再構成し、その場面で可能な物体の組み合わせタスクをサンプリングする。各タスクについて、幾何学的な再構成上で純粋に運動学を使って把持と軌道を計画し、対象環境に視覚的に一致する写真のようにリアルな実演を描画する。物理的な動力学を処理に取り込むのは、接触力の相互作用が結果を決める把持形成時だけであり、相互作用データセットで一度事前学習した接触モデルを用いる。 合成実演の下流での有用性を評価するため、スキャンから導入までの一貫した作業手順を二つの設定で実装した。一つは再現性を確保するために実世界の代わりとしたシミュレーション場面、もう一つは実機のUR10eロボットを備える実世界の作業空間である。各設定で、合成実演だけで学習した標準的な拡散方策は、それぞれ95.1%と84.2%の成功率を達成した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Training a visuomotor policy calls for abundant demonstrations that closely match the target environment, yet collecting them anew remains expensive. Existing demonstration synthesis methods reduce this cost but remain constrained by high manual effort, limited visual fidelity, or heavy reliance on physics simulators. This paper introduces GaussianFactory, a high-fidelity data engine that mass-produces demonstrations with a single video scan as its only human input and no physics engine in the generation loop. Specifically, GaussianFactory reconstructs the scene as an editable 3D Gaussian Splatting (3DGS) replica and samples from the object-combination tasks the scene affords. For each task, it plans grasps and trajectories purely kinematically on the geometric reconstruction, rendering photorealistic demonstrations that visually match the target environment. Physical dynamics enter the pipeline only where contact force interactions dictate the outcome---during grasp formation, via a learned contact model pretrained once on an interaction dataset. To evaluate the downstream utility of the synthesized demonstrations, we implement an end-to-end scan-to-deployment workflow in two setups: a simulated scene that stands in for the real world to enable reproducibility, and a real-world workspace with a physical UR10e robot. In each setup, a standard diffusion policy trained solely on the synthesized demonstrations achieves 95.1% and 84.2% success rates, respectively.
arXiv ID: 2609.21112 / 要約の誤りについて