動く物体を長時間操作するロボットの評価基盤
MotionForge: A Data Generation Pipeline and Large-Scale Benchmark for Long-Horizon Manipulation of Dynamic Objects with Domain Shifts
この論文をやさしく読む
ひとことで言うと
動き続ける物体をロボットが長時間操作する際の性能を、環境条件の変化も含めて調べるシミュレーション基盤。
何に役立つ?
動的な操作方策の比較や、背景・照明・速度などが変わる場合の弱点の発見に役立つ。
この研究の面白いところ
方策の推論中も環境が進む実行手順と、単一要因・複数要因の変化を分けた評価を用意した。
どこまで分かった?
シミュレーション上の40課題での評価であり、実ロボットでの性能は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
学習に基づくロボット方策は進歩しているが、評価の中心は静的またはほぼ静的な環境である。動的な操作では、ロボットが観測、判断、動作する間も物体や場面が変化し続ける。近年の動的シミュレーションの評価基盤は、単純な動きに対する短時間の反応的なやり取りに偏り、環境条件が変わった場合の体系的な評価や、モデルに依存しないリアルタイムの実行手順への対応が限られる。本稿は、動的操作における環境条件の変化と長時間の相互作用を併せて評価するための、大規模なシミュレーション評価基盤とデータ生成手順MotionForgeを提示する。11種類の動き方にまたがる40の動的な相互作用課題を含み、そのうち17課題は長時間の操作に対応する。主な新規性は2つある。第一に、背景だけの変化など単一要因の変化と、物体・背景・照明・速度の同時変化の両方に対し、方策の頑健性を測る体系的な評価手順である。第二に、環境が方策の推論時間とは独立に変化し続ける、遅延を考慮した分離型の実行手順である。代表的な汎用ロボット方策を広く評価したところ、複数要因が同時に変わる場合に大きな限界が見られた。これらの結果は、現行の方策能力と、条件変化の下で動く物体を頑健に長時間操作するための要件との隔たりを示し、MotionForgeを将来の身体性AI研究の評価環境として位置付ける。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent advances in learning-based robot policies have demonstrated promising progress, yet they are predom- inantly evaluated in static or quasi-static environments. In dynamic manipulation, objects and scenes continuously evolve while the robot perceives, reasons, and acts. However, recent dynamic simulation benchmarks largely focus on short-horizon, reactive interactions with simple motion patterns and offer limited support for both systematic evaluation under domain shifts and model-agnostic real-time execution protocols. To bridge these gaps, we introduce MotionForge, the first large- scale simulation benchmark and data-generation pipeline tailored to jointly evaluate domain shifts and long-horizon interaction in dynamic manipulation. MotionForge comprises 40 dynamic interaction tasks spanning 11 distinct motion patterns, with dedicated support for 17 long-horizon tasks. Our benchmark introduces two key novelties: (1) a systematic evaluation protocol for assessing policy robustness under both single-factor (e.g., only backgrounds shift) and joint domain shifts (e.g., simultaneous shifts of objects, backgrounds, lighting, and speed); and (2) a decoupled, latency-aware execution protocol where the environ- ment continuously evolves independently of policy inference time. Extensive evaluations of representative general-purpose robot policies on our benchmark reveal substantial limitations under joint domain shifts. These findings expose a critical gap between current policy capabilities and the requirements of robust long- horizon manipulation of dynamic objects under domain shifts, establishing MotionForge as a comprehensive testbed for future research in embodied AI.
著者のコメント
9 pages
arXiv ID: 2609.25689 / 要約の誤りについて