自動運転の衝突場面を二つの言語モデルで探索
Teach-to-Crash: A Closed-Loop Student-Teacher LLM Framework for Collision-Inducing Test Scenario Generation
この論文をやさしく読む
ひとことで言うと
二つの言語モデルで自動運転の試験場面を探索し、衝突を起こす場面の生成を改善する研究。
何に役立つ?
考えられる用途は、シミュレーション上で自動運転の失敗を探し、後の分析に使う場面を集めること。
この研究の面白いところ
探索が停滞したときだけTeacherが介入する。衝突率だけでなく多様性と回避可能性の代理指標も評価した。
どこまで分かった?
結果は速度方策を変えた二設定のCARLA事例研究に限られる。実走行での安全性向上を示したものではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
シミュレーションで自動運転システム(ADS)を検証するには、実行可能で多様かつ後の失敗分析に役立つ場面を生成しながら、まれだが安全上重要な失敗を発見できる試験の仕組みが必要である。本研究は、制約された自車中心の場面表現、探索の停滞を検知する制御、二つの言語モデルを用いる構成を組み合わせた、閉ループの試験枠組みTeach-to-Crashを導入する。推論能力の高いTeacherモデルが適応的な探索制御を担い、推論量の少ないStudentモデルが厳密なJSON形式でシミュレーター実行可能な場面を出力する。Teacherは、移動窓で見た衝突率と衝突までの時間の指標が停滞したときだけ介入し、探索の方向を変えるための戦略的な指導を与える。自車の速度方策を変えた二つの実験設定を含むCARLAの事例研究で、Teach-to-Crashは比較手法中最高の衝突ヒット率90.79%、最短の平均衝突までの時間18.31秒、競争力のある衝突発見率136.21を達成した。PAFOTの平均衝突発見率は179.44と高いが、ばらつきはかなり大きかった。Teach-to-Crashは多様性でも最高値0.547を示し、両設定のCARLA Traffic Manager制御器で平均すると、回避可能性に基づく有用性の代理指標も比較手法中最高の60.04%だった。評価したCARLAの範囲では、閉ループでの二つの言語モデルによる推論が、制約された実行可能なプログラム空間で敵対的なシミュレーション試験を導き、頻繁で構造的に多様、かつ回避可能と評価される割合がより高い失敗を生成できることを示す証拠となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Validating Autonomous Driving Systems (ADS) in simulation requires testing architectures that can discover rare, safety-critical failures while generating scenarios that are executable, diverse, and useful for downstream failure analysis. We introduce Teach-to-Crash, a closed-loop testing framework that combines a constrained ego-centric scenario representation, stagnation-aware search control, and a dual-LLM architecture for adaptive failure discovery. A high-reasoning Teacher LLM acts as an adaptive search controller, while a low-reasoning Student LLM emits simulator-executable scenarios in a strict JSON schema. The Teacher intervenes only when rolling collision rate and time-to-collision metrics stagnate, providing strategic guidance to redirect the search. In a CARLA case study with two experimental setups that vary the ego vehicle's speed policy, Teach-to-Crash achieves the highest Collision Hit Rate (90.79%), the shortest mean Time-to-Collision (18.31 s), and a competitive Collision Discovery Rate (136.21). PAFOT attains a higher mean CDR (179.44), but with substantially larger variance. Teach-to-Crash also yields the highest diversity (0.547) and, averaged across both setups on the CARLA Traffic Manager controller, the highest avoidability-based usefulness proxy (60.04%) among the compared methods. These results, within the evaluated CARLA scope, provide evidence that closed-loop dual-LLM reasoning can steer adversarial simulation-based testing over a constrained executable program space, generating failures that are frequent, structurally diverse, and assessed as more frequently avoidable.
著者のコメント
41st IEEE/ACM International Conference on Automated Software Engineering (ASE) AgenticDev (2026)
arXiv ID: 2609.27296 / 要約の誤りについて