arXiv論文メモ
新着一覧
cs.RO / cs.LG · 査読状況未確認

構造化されていない道路へ強化学習の運転を移す

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Thomas Steinecker, Denis Trescher, Alexander Bienemann, Thorsten Luettel, Mirko Maehlisch

この論文をやさしく読む

ひとことで言うと

シミュレーションと実車で共通の鳥瞰表現を使い、未舗装区間などを走る強化学習モデルを実車へ移します。

何に役立つ?

考えられる用途は、構造化が弱い走行環境における学習済み運転方策の実車導入です。

この研究の面白いところ

認識表現をそろえるだけでなく、行動を軌道として整合させます。異なる2車両で合計17.3 kmを人の介入なしに走行しました。

どこまで分かった?

実証は3.0 kmの試験コースで最高33.6 km/hの条件を含みます。一般公道やあらゆる交通状況での性能を示す評価ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

強化学習は、人を超える性能や自律的に学習するポリシーの可能性を持つ有望な方法である。しかし、構造化されていない環境でシミュレーションから実世界へ移すことの難しさから、実世界の自動運転、特にそのような環境での応用はまだ少ない。 本研究では、追加学習なしでシミュレーションから実世界へ転移する、端から端までのポリシーの枠組みMILERを提示する。オフライン訓練では、独自の意味的な中間レベル表現(MLR)シミュレーターを用い、強化学習でポリシーネットワークを訓練する。その制御出力は、車両の二輪モデルへ直接適用する。 実車への導入時には、カメラとLiDARのデータをBEVFusionで処理し、MLRシミュレーターと整合する意味的な鳥瞰表現を生成する。ポリシーネットワークが生成した行動は、そのまま実車に適用しない。代わりに軌道整合戦略を用い、知覚と制御の両方についてゼロショット転移を可能にする。 各種の障害物、ヘアピンカーブ、最高33.6 km/hの速度、未舗装区間など、多くの課題を含む多様な試験コースで広く評価した。長さ3.0 kmの試験コースで、異なる2台の車両を用いて合計17.3 kmを、人の介入なしで走行し、手法の有効性を示した。さらに、ソフトウェア全体はJetson AGX Orin上で動作する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Reinforcement learning constitutes a promising approach owing to its potential for superhuman performance and self-learned policies. However, its application to real-world autonomous driving remains scarce, particularly in unstructured environments, because of the challenges associated with sim-to-real transfer for unstructured environments. In this work, we present MILER, an end-to-end policy framework with zero-shot sim-to-real transfer. During offline training, we employ a custom semantic mid-level representation (MLR) simulator and train the policy network using reinforcement learning, with its control outputs applied directly to a bicycle model. During deployment on the real vehicle, camera and LiDAR data are processed by BEVFusion to generate a semantic bird's-eye-view representation consistent with that of the MLR simulator. The actions generated by the policy network are not applied directly to the real vehicle. Instead, we employ a trajectory-alignment strategy that enables zero-shot sim-to-real transfer of both perception and control. We extensively evaluate the proposed framework on a diverse test track comprising numerous challenges, including various obstacles, hairpin curves, velocities of up to 33.6 km/h, and off-road sections. In total, we drove 17.3 km with two different vehicles on a 3.0 km test track without human intervention, thereby demonstrating the effectiveness of our approach. Furthermore, the entire software stack runs on a Jetson AGX Orin.

著者のコメント

Evaluation video: https://www.youtube.com/watch?v=IZli3Z87URI

arXiv ID: 2609.20747 / 要約の誤りについて