arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

スパイキング神経網でドローンの狭い開口部通過を学習

Spiking Neural Network Actor-Critic Proximal Policy Optimization Control for Autonomous UAV Navigation Through Constrained Openings in Civil Infrastructure and Buildings

Francis Noah Walugembe, Maciej Wielgosz, Tomaž Goričan, Matej Mertik

この論文をやさしく読む

ひとことで言うと

ドローンが狭い窓状の開口部を順に通過する動作を、スパイキング神経網と強化学習で学ばせています。

何に役立つ?

考えられる用途は、橋やトンネル、建物内などの点検での自律航行です。計算コストを課題として、別の神経網構成を検討しています。

この研究の面白いところ

全体の成功率63.77%と、後期段階の90%超という成績を分けて示し、通過窓数も評価しています。

どこまで分かった?

要旨は実機飛行かシミュレーションかを明示せず、計算量や消費電力の削減値も示していません。また「3,000回超」と1,913回・63.77%の分母には説明不足があり、数値は原文どおり保持しています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

制約のある3次元環境における無人航空機の自律航行は、ロボティクス分野の課題となってきた。土木インフラ点検での自律無人航空機の利用には、橋梁、トンネル、構造物の点検が含まれる。深層強化学習を無人航空機の自律航行に適用することで、制約のある環境でも成果が得られている。しかし、そのアルゴリズムの計算コストが、自律航行への適用を制限している。 本論文では、制約が連続する環境で無人航空機を自律航行させるため、スパイキングニューラルネットワークに基づく近接方策最適化(PPO)アルゴリズムの利用を提案する。提案法は、スパイクに基づくアクター・クリティック強化学習をPPOと統合し、確率的なガウス方策を用いて無人航空機を自律航行させる。制約のある3次元環境で実装した結果、3,000回を超えるエピソードのうち1,913回を完了し、1エピソード当たり平均2.10個の窓を通過した。成功率は63.77%であり、アルゴリズムの後期段階では90%を超える成功率を達成した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Autonomous navigation of unmanned aerial vehicles in constrained three-dimensional environments has been a challenge in the robotics domain. The application of autonomous unmanned aerial vehicles in civil infrastructure inspection involves the use of such vehicles in bridge inspection, tunnel inspection, and structural inspection. The use of deep reinforcement learning in the autonomous navigation of unmanned aerial vehicles has been successful in constrained environments. However, the computational cost of the algorithm limits the application of the algorithm in the autonomous navigation of unmanned aerial vehicles. This paper proposes the use of the spiking neural network-based Proximal Policy Optimization algorithm in the autonomous navigation of unmanned aerial vehicles in constrained sequential environments. The proposed algorithm integrates the use of spike-based actor-critic reinforcement learning with the Proximal Policy Optimization algorithm. The proposed algorithm uses the stochastic Gaussian policy in the autonomous navigation of unmanned aerial vehicles. The proposed algorithm was implemented in the autonomous navigation of unmanned aerial vehicles in constrained 3D environments. The proposed algorithm was successful in completing 1913 episodes out of more than 3000. The proposed algorithm was successful in passing an average of 2.10 windows per episode. The proposed algorithm was successful in achieving a success rate of 63.77%. The proposed algorithm was successful in achieving success rates of more than 90% in the later stages of the algorithm.

arXiv ID: 2609.23643 / 要約の誤りについて