arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

行動を量子化したオフライン強化学習で自動駐車を学ぶ

Learning Reliable Parking Policies via Offline Reinforcement Learning with Quantized Action Representations

Zewei Yang, Zengqi Peng, Jun Ma

この論文をやさしく読む

ひとことで言うと

事前に集めた駐車動作から学び、周囲の車との相互作用も考えて経路点を選ぶ自動駐車の方法です。

何に役立つ?

混雑した駐車空間での動作生成を研究するための学習枠組みになります。連続した経路を離散的な行動記号に変え、データに乏しい動作の価値を過大評価しないようにします。

この研究の面白いところ

障害物のLiDAR特徴を目標姿勢に応じて調整し、状態依存のトークン化と保守的Q学習を組み合わせています。対話的な場面と非対話的な場面の双方をデータに含めています。

どこまで分かった?

CARLAシミュレーターの閉ループ実験で、比較対象中の最高成功率と未見の駐車区画への転移を報告しています。要旨には具体的な成功率や実車検証は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

駐車は、都市環境で動く自動運転車にとって日常的でありながら、安全上重要な課題である。しかし、障害物が多く構造の明確でない駐車空間に、周囲の車両との相互作用の不確実性が加わることで、信頼できる操車動作の生成が難しくなる。これに対し、相互作用を考慮した自動駐車のため、経由点単位のオフライン強化学習の枠組みを開発する。 具体的には、階層型の熟練方策による走行から、経由点を回転させるデータ拡張を用いて専用の駐車データセットを構築し、他車との相互作用があるシナリオとないシナリオの両方を含める。方策はコンパクトな状態表現を条件とし、その中でLiDARに基づく障害物特徴を、特徴ごとの線形変調によって目標姿勢に適応させる。さらに、状態を条件とするトークナイザで連続的な経由点列を離散的な行動トークンへ量子化する。その上で保守的Q学習を行い、データによる裏づけが乏しい行動の価値の過大評価を抑える。 高忠実度のCARLAシミュレータで広範な閉ループ実験を行った。提案手法はすべての比較手法の中で最も高い駐車成功率を達成し、未見の駐車区画にも安定して転移した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Parking is a routine yet safety-critical task for autonomous vehicles operating in urban environments. However, cluttered and weakly structured parking spaces, compounded by the interactive uncertainty from surrounding vehicles, hinder reliable maneuver generation. To address these challenges, we develop a waypoint-level offline reinforcement learning framework for interaction-aware autonomous parking. Specifically, a dedicated parking dataset is constructed from hierarchical expert rollouts with rotational waypoint augmentation, covering both non-interactive scenarios and interactive ones. The policy is then conditioned on a compact state representation, in which LiDAR-based obstacle features are adapted to the target pose via feature-wise linear modulation. A state-conditioned tokenizer further quantizes continuous waypoint sequences into discrete action tokens, over which conservative Q-learning is performed to suppress value overestimation on poorly supported actions. Extensive closed-loop experiments are conducted in the high-fidelity CARLA simulator. The proposed framework attains the highest parking success rate among all baselines and transfers reliably to unseen parking slots.

arXiv ID: 2609.19894 / 要約の誤りについて