変化する課題への適応と忘却を神経進化で比較する
Continual Reinforcement Learning with Neuroevolution
この論文をやさしく読む
ひとことで言うと
課題が変わり続けるとき、勾配による学習の代わりに、ネットワークの集団を変異・選択させる方法を比較しています。進化戦略は、新しい課題への適応と以前の能力の保持を両立しやすい結果でした。
何に役立つ?
変化する課題を学び続けるシステムの訓練方法を選ぶ材料になります。重みに揺らぎを与えても解ける領域の広さを見ることが、適応と忘却を評価する手掛かりになっています。
この研究の面白いところ
性能比較に加え、解の周りの重み空間を調べて手法の違いを説明しています。行動の多様性を強めると適応は進む一方、忘却も増える関係が示されています。
どこまで分かった?
近傍の重なりと安定性・可塑性の関係は相関として述べられています。要旨には個別の環境名、比較数値、計算コストは示されず、重み摂動の一般的な有効性は示唆として位置付けられています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
タスクが継続的に変化する強化学習(RL)では、可塑性が失われる原因と対策について多くの研究があるが、適応と忘却の良好なバランスを一貫して達成するRL手法はまだない。本研究では、別の最適化の枠組みである神経進化(NE)に着目する。これは、ニューラルネットワークの集団に突然変異と選択を適用し、重み空間を直接探索するアルゴリズムである。数百パラメーターから百万パラメーター規模の方策を用い、幅広い環境と環境変化において、進化戦略(ES)および遺伝的アルゴリズム(GA)を、最先端の継続RLの各手法と集団ベースRLと比較する。ESは安定性と可塑性の良好な釣り合いを最も一貫して実現し、GAは最も可塑的である一方、ESより多くを忘れる。 これを説明するため、各手法で得られた解の周囲の収益地形を調べる。ESは最も広い近傍、すなわち方策に摂動を加えてもタスクを解ける重み空間の領域を見いだす。連続するタスクの近傍同士の重なりの大きさは、各手法の安定性と可塑性の釣り合いと相関する。新奇性探索によってGAの行動の多様性に報酬を与えると、集団の可塑性はさらに高まるが、忘却が増える。最後に、RLでよく報告される可塑性喪失の兆候は、NEにはそのまま当てはまらない。これらの結果は、継続的なタスク変化の下でNEがRLに対抗できる選択肢であることを示し、重み空間で摂動を加えながら訓練することが、より広い継続学習に有用な仕組みとなり得ることを示唆する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method's solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method's stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.
arXiv ID: 2610.01583 / 要約の誤りについて