水中ロボットの故障診断と復旧を言語モデルで検証する基盤
A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies
この論文をやさしく読む
ひとことで言うと
水中機の異常時だけ言語モデルに診断と復旧を頼む構成を、多数の故障シミュレーションで評価します。
何に役立つ?
通信が難しい水中での自律回復を検討するため、単発の成功例に頼らず比較する実験基盤になります。
この研究の面白いところ
480試行で質量移動故障を調べ、上位三仮説に原因を含む率は先端モデル85〜90%、最良ローカルモデル60〜78%でした。
どこまで分かった?
一種類の故障を中心としたシミュレーションです。診断の良さと運用判断の良さはこのデータでは結びつかず、診断率を任務回復率とみなせません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
信頼できる通信が届かない場所で動作する自律型無人潜水機(AUV)は、人の介入なしに故障から復旧しなければならない。本研究では、通常の運用を従来の決定論的な階層制御による自律機能が担い、機上の異常検出が想定範囲外の性能を検知した際には、呼び出し可能な大規模言語モデル(LLM)が診断と復旧計画を担当する構成を調べる。言語モデルは確率的であるため、厳密な評価には個別の実演ではなく、多数の試行をまとめたアンサンブル試験が必要になる。 Cで実装したリアルタイムの機体ソフトウェアと、物理モデルに基づく故障注入、構造化プロンプト、言語モデルとのやり取り、ミッションファイルの生成・検証・実行、LLM判定器による採点を担う上位の統括層を結合した、閉ループのシミュレーション構成を提示する。SPAR(Simulation Platform for AUV Recovery)と呼ぶこの枠組みは、故障の具体的な実現例、プロンプト構造、推論モデル、ミッション条件を変えた評価に対応する。質量移動故障について、これらを変化させた480回のSPAR試行を行い、最先端モデル一つと、既製のローカル配備可能なLLM三つを評価した。 診断を最も大きく左右するのはモデルの選択である。最先端モデルでは試行の85〜90%で重心(CG)移動の機構が上位三つの仮説に入ったのに対し、最良のローカルモデルでは60〜78%だった。推論の分析は、ローカルモデルの成功が診断手順を最後まで実行することと関連し、弱いモデルでは、アクチュエーターが指令に追従しているにもかかわらず、昇降舵の故障という結論に早まって固執することが多いと示している。このデータセットでは、診断性能と運用上の意思決定性能には結びつきが見られない。 本研究の貢献は、予期しない故障への復旧を検出から影響軽減まで拡張する構成と、低消費電力AUVにおけるLLM支援型ミッション管理を評価するアンサンブル手法である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.
著者のコメント
6 pages, 3 figures, 2 tables. Accepted for presentation at the 2026 IEEE/OES Autonomous Underwater Vehicles Symposium (AUV 2026), Southampton, UK. This is the author-accepted manuscript
arXiv ID: 2609.20620 / 要約の誤りについて