arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

未舗装路での強化学習を速める地形考慮の世界モデル

Verti-WM: A Physics-Aided Exteroceptive World Model for Off-Road Reinforcement Learning

Chenhui Pan, Tong Xu, Xuesu Xiao

この論文をやさしく読む

ひとことで言うと

地形情報と物理モデルを使い、未舗装路を走る車両の学習を速める研究。

何に役立つ?

高忠実度シミュレーションを繰り返す費用を抑えた方策学習の設計に役立つ。

この研究の面白いところ

硬い地面と変形する地面を別のモデルで扱い、実機でも比較した点。

どこまで分かった?

実機の成功率は指定されたVerti-4-Wheelerでの結果であり、他の車両や地形への一般化は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

未舗装路を走るための強化学習には車両と地形の相互作用データが大量に必要だが、高忠実度のシミュレーションで集めるには費用がかかる。世界モデルは方策の最適化中にシミュレータでの走行を置き換える有望な方法だが、未舗装路では車体内部の観測だけでは足りず、外部から得る地形情報に応じて状態遷移を予測する必要がある。硬い地面と変形する地面の両方を扱う必要もあり、データに基づく方法と物理に基づく方法にはそれぞれ強みがある。本研究は、固定したTransformerで硬い地形を扱い、神経記号的な地盤力学モデルで変形する地形を扱って、両者を再帰的に融合する物理補助型の外部観測世界モデルVerti-WMを提案する。予測した位置ごとに与えられた地図から標高と意味情報を取得して融合を条件付け、高忠実度シミュレータを再度使わずに方策最適化用の6自由度の走行予測を行う。予測誤差はデータ駆動の基準法に比べ34.6%、物理ベースの基準法に比べ21.7%低下した。Verti-WMだけで学習した方策は、高忠実度シミュレータで直接学習する場合と同程度の課題成功率を得て、計算時間は23.6分の1になった。実環境のデータでも検証し、学習した実環境の運動特性の中で方策を最適化した結果、Verti-4-Wheeler上で成功率80%を達成した。シミュレーションから実機へ直接移した場合は40%だった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Reinforcement learning for off-road navigation requires extensive vehicle-terrain interaction data, which are costly to collect in high-fidelity simulation. World models offer a promising alternative by replacing simulator roll-outs during policy optimization. However, an off-road world model must condition state transitions on exteroceptive terrain information, which proprioception alone does not provide. This challenge is further amplified by the need to model both rigid and deformable terrain, where data-driven and physics-based approaches offer complementary strengths. We propose Verti-WM, a physics-aided exteroceptive world model that recurrently fuses a frozen Transformer for rigid terrain and a neuro-symbolic terramechanics model for deformable terrain. Elevation and semantic observations queried from a supplied map at each predicted pose condition fusion, enabling six-degree-of-freedom rollouts for policy optimization without further simulator access. Verti-WM reduces prediction error by 34.6% and 21.7% over data-driven and physics-based baselines, respectively. Policies trained entirely within Verti-WM achieve comparable task success rates while reducing computation time by 23.6X relative to direct training in the high-fidelity simulator. We further validate Verti-WM using real-world data, enabling policy optimization within learned real-world kinodynamics and achieving a 80% success rate on the Verti-4-Wheeler platform, compared with 40% for direct sim-to-real transfer.

arXiv ID: 2609.23118 / 要約の誤りについて