画像から変形物体のロボット用シミュレーション素材を作る
DiagGen: Agentic Generation of Deformable Assets with Sim-based Diagnostics for Robotic Simulation
この論文をやさしく読む
ひとことで言うと
1枚の画像から変形する物体のシミュレーション用モデルを作り、物理シミュレーションで見つけた問題を生成工程に戻して直す手法。
何に役立つ?
考えられる用途は、ロボットの把持や移動の計画を仮想環境で試すための素材作成である。論文では、生成物を物理シミュレーターで使えることを示している。
この研究の面白いところ
物体の見た目だけで終わらず、シミュレーターで領域ごとの反応を調べ、どの生成段階を直すかに結び付ける。
どこまで分かった?
40個の素材で品質の一定の改善を報告している。改善の幅の具体的な数値や、あらゆる物体への適用は要旨には示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ロボットが変形する物体を操作する課題には、シミュレーションで使える素材が欠かせない。しかし従来の生成手法は、物理的なもっともらしさを生成後に評価することが多く、シミュレーション中の物体の反応を、先行する生成段階の誤りを修正するためのフィードバックに利用していない。本研究は、実環境で撮影された1枚の画像から、生成、シミュレーション、診断、改良を繰り返して、シミュレーション可能な変形物体の素材を作るDiagGenを提案する。部品を考慮した形状と材料パラメータを構築し、視覚言語モデルに基づくエージェントが意味上重要な領域を選び、物理シミュレーターで調べ、材料の反応を観察して、証拠に基づく修正の手がかりを担当する生成段階に振り分ける。40個の素材を用いた実験では、診断が有用な修正の手がかりを与え、生成された変形物体の品質をある程度改善できることを示した。また、視覚基盤モデルで生成した素材にはシミュレーションできないものがある一方、DiagGenで生成した変形物体は高精度な物理シミュレーターに直接投入でき、接触の多いピックアンドプレース作業の計画とシミュレーションに利用できることを示した。プロジェクトのウェブサイトは https://diaggen.github.io/ である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
While simulation-ready deformable assets are essential for in-silico robotic manipulation tasks, existing generation frameworks typically assess physical plausibility after generation, leaving an object's simulated response unused as feedback for repairing upstream errors. We present DiagGen, an agentic framework that turns a single in-the-wild image into a simulation-ready deformable asset through a generate--simulate--diagnose--refine loop. DiagGen constructs part-aware geometry and material parameters, then uses a VLM (vision-language model)-based agent to select semantically informative regions, probe them in a physics simulator, observe material responses, and route evidence-backed repair cues to the responsible generation stage. Experiments on 40 assets show that diagnostics provides useful repair cues and can moderately improve the quality of generated deformable assets. Finally, we show that unlike assets generated from visual foundation models which may not be simulatable, DiagGen-generated deformables can be directly dropped into a high-fidelity physical simulator for the planning and simulation of contact-rich pick-and-place tasks. The project's website is https://diaggen.github.io/.
arXiv ID: 2609.23103 / 要約の誤りについて