実機性能に近いロボット操作のシミュレーション評価
X2Real: an eXtensive simulation benchmark for real-world generalist policies
この論文をやさしく読む
ひとことで言うと
ロボット操作のシミュレーション評価を、実機の結果と近づけ、幅広い課題で公平に比較できるようにしたベンチマーク。
何に役立つ?
汎用ロボット方策を実機へ持ち込む前に比較する際、課題の幅や訓練・評価の分離をそろえる基盤として使える可能性がある。実機性能を完全に予測できるという結果ではない。
この研究の面白いところ
見た目と物理条件を実機へ合わせ、評価結果の線形相関0.84を報告した。10の能力軸、44の長期課題、約300時間の注釈付き軌道を組み合わせている。
どこまで分かった?
相関0.84は要旨に記された校正と評価での結果である。未知のロボット機体や作業条件にも同じ相関が成り立つかは示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
多目的に使えるロボットの物体操作方策は急速に発展しているが、信頼できる評価は難しい。既存のシミュレーションによるベンチマークには、実機との大きな差、狭い課題範囲、訓練と評価の手順の曖昧さによる不公平な比較という基本的な問題がある。従来の研究はこれらを部分的にしか解決しておらず、忠実性、多様性、公平性を同時に備えない。また、固定されたベンチマークの設計は、方策の長期的な改良を支えにくい。 著者らは、Nvidia Isaac Lab-Arenaに基づき、実世界でのロボット操作方策の性能を忠実に評価するため、拡張可能なシミュレーションベンチマークX2Realを提示する。忠実性、多様性、公平性の3原則に従い、シミュレーションの見た目と物理特性を実機へ合わせた結果、シミュレーションと実機での評価結果の線形相関は0.84になった。基本操作から、視覚による対象の特定、言語理解、双腕制御まで、10の能力軸と階層的な長期課題44件を含む。さらに、複数軸にわたる条件のランダム化と、厳密に分離された訓練・評価手順によって、ベンチマークに過度に合わせた解法を抑え、評価の信頼性を高める。専用の物理的な領域特化言語を用いたManaシミュレーション環境は、課題の部品化と性能の反復分析を支え、注釈付きのシミュレーション軌道データを約300時間分備える。X2Realは、現実との評価差を縮め、汎用ロボット操作方策の発展を支える、進化させられる評価基盤である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Generalist robot manipulation policies have developed rapidly, yet their reliable evaluation remains challenging due to fundamental flaws in existing simulation benchmarks: prominent sim-to-real gaps, narrow task coverage, and unfair evaluation caused by ambiguous training-test pipelines. Prior works only partially resolve these issues and lack simultaneous faithfulness, diversity, and fairness, while static benchmark designs fail to sustain long-term policy development. We present X2Real, an evolvable simulation benchmark for faithfully evaluating the real-world performance of robotic manipulation policies based on Nvidia Isaac Lab-Arena. Following three core principles (faithfulness, diversity, and fairness), X2Real calibrates simulation visual and physical properties to align with real hardware, achieving a 0.84 linear correlation between simulated and real-robot evaluation results. It features a comprehensive taxonomy with 10 capability dimensions and 44 hierarchical long-horizon tasks, covering basic manipulation skills and advanced capacities such as visual grounding, language understanding, and bimanual control. We further adopt multi-axis domain randomization and strictly disjoint training-evaluation pipelines to mitigate benchmark exploitation and ensure credible evaluation. Powered by a custom physical domain-specific language, the Mana simulation ecosystem supports modular task design and iterative performance analysis, alongside a nearly 300-hour annotated simulation trajectory dataset. X2Real offers a faithful, diverse, and fair evolving evaluation infrastructure, effectively bridging the sim-to-real evaluation gap and supporting the advancement of generalist robotic manipulation policies.
arXiv ID: 2609.27449 / 要約の誤りについて