到達可能範囲を使う強化学習で地球・火星間の軌道を設計
Reachability-Informed Reinforcement Learning for Multi-Impulse Interplanetary Transfers
この論文をやさしく読む
ひとことで言うと
物理的に到達できる範囲から強化学習で経由点を選び、古典的な軌道計算で実際の移動を組み立てる方法です。
何に役立つ?
考えられる用途は、出発条件が少しずつ異なる惑星間飛行の軌道を、毎回学習し直さず設計することです。実証は二体問題の地球・火星間の数値ベンチマークです。
この研究の面白いところ
強化学習に軌道変更を直接すべて決めさせず、経由点の選択を担当させています。複数状態での学習により、コストを少し増やして実行可能な出発条件を大きく広げました。
どこまで分かった?
10,000件すべての成功は、固定の目標状態・遷移時間と、試した出発条件での数値結果です。実宇宙機の飛行試験ではなく、任意の摂動やすべての惑星間遷移での保証ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
強化学習は、宇宙機の軌道設計に再利用可能な逐次意思決定機構をもたらす可能性があり、学習による判断と基礎となる軌道変更の幾何を結び付ける方策インターフェースが求められる。本論文では、決定論的な複数インパルスによる惑星間遷移に向けた、到達可能性解析を組み込む強化学習(RARL)を開発する。中間経由点の選択を学習による意思決定の中心に据える。局所的な1次の到達可能性は、有界な速度摂動を次の節点位置の楕円体集合へ写し、方策はその中から経由点を選ぶ。次にLambert問題による再構成が、力学的に整合した弾道弧に沿って選ばれた経由点へ到達するための軌道変更を決め、学習による遷移幾何の選択を古典的な宇宙航行力学と結び付ける。終端の2インパルス再構成によってランデブーを完了する。その際、線形の軌道変更要求量評価を報酬整形に用いる。 数値研究では、二体問題の地球・火星ベンチマーク上でこのインターフェースの性質を調べる。独立した3回の学習で、RARLの平均軌道変更コストは10.23 km/sとなり、検証済みの局所的な逐次凸計画法による基準値を1.72%上回った。ばらつきのある初期状態で学習すると、目標状態と遷移時間を固定した出発条件の族にわたり、方策を再利用できる範囲が広がる。独立に学習した3つの複数状態方策は、それぞれ学習から除外したモンテカルロ出発条件10,000件すべてを、インパルス上限違反なしで完了した。単一状態方策の平均実行可能率は6.49%だった。標本上の実行可能性がこのように広がる一方、平均公称軌道変更コストは0.61%増加したが、各出発条件での追加学習は不要だった。これらの結果は、到達可能性を組み込んだ意思決定インターフェースが、ベンチマークに匹敵する質の軌道構築と、ばらついた出発条件間での方策再利用を支えることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Reinforcement learning offers the prospect of a reusable sequential decision-making mechanism for spacecraft trajectory design, motivating policy interfaces that connect learned decisions to the underlying maneuver geometry. This paper develops Reachability Analysis-Informed Reinforcement Learning (RARL) for deterministic multi-impulse interplanetary transfers, placing intermediate waypoint selection at the center of the learned decision process. Local first-order reachability maps bounded velocity perturbations into an ellipsoidal set of next-node positions, within which the policy selects its waypoint. Lambert reconstruction then determines the corresponding maneuver to reach this selected waypoint along a dynamically consistent ballistic arc, coupling learned transfer-geometry selection with classical astrodynamics. A terminal two-impulse reconstruction completes the rendezvous, supported by a linear maneuver-demand assessment used for reward shaping. Numerical studies characterize this interface on a two-body Earth-Mars benchmark. Across three independent training runs, RARL achieves a mean maneuver cost of 10.23 km/s, 1.72% above a validated local sequential convex programming reference. Training over dispersed initial states extends policy reuse across a departure family with fixed target state and transfer duration. Each of the three independently trained multi-state policies completes all 10,000 held-out Monte Carlo departures without impulse-cap violations, compared with a mean feasibility rate of 6.49% for single-state policies. This broader sampled feasibility is accompanied by a 0.61% increase in mean nominal maneuver cost, without further training across departures. These results demonstrate that a reachability-informed decision interface supports benchmark-quality trajectory construction and policy reuse across dispersed departure conditions.
著者のコメント
Preprint. 23 pages, 10 figures
arXiv ID: 2610.01344 / 要約の誤りについて