arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

過去の計算結果を再利用する組合せ最適化の学習

MiLoop: Selective Memory Propagation for Neural Combinatorial Optimization

Changliang Zhou, Yuanyao Chen, Rongsheng Chen, Zhiyun Lin, Zhenkun Wang

この論文をやさしく読む

ひとことで言うと

最適化の解を一手ずつ作る際に、前の段階で計算した表現を記憶として引き継ぐ学習手法です。

何に役立つ?

正解の解を学習用に準備しなくても、解の品質を報酬にして構築方策を学ぶために役立ちます。異なる規模の問題への適用を実験しています。

この研究の面白いところ

複数段階の計算を単なる手順として使うだけでなく、記憶を更新する機会として利用し、浅い方策でも動的な表現を作れるようにしています。

どこまで分かった?

4種類の問題と100〜1000万ノードの範囲での計算実験です。要旨には各問題の名称、具体的な最適性ギャップ、実行時間の値はなく、常に最適解を得る保証は述べられていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

構築型ニューラル組合せ最適化(NCO)は、組合せ最適化問題(COP)の解を段階的に構築することを学ぶ有望な枠組みとして登場しており、人手で設計した規則への依存を減らし、高速な推論を可能にする。動的埋め込みを使う多くの手法は良好な汎化性能を示すが、通常は各段階で深い注意機構の積み重ねを使って部分問題の表現を一から作り直す。この種の高性能な手法の多くは、効率よく学習するために正解の解や擬似的な正解ラベルに依存するか、強化学習(RL)中に探索空間を積極的に枝刈りする。 これらの制約に対処するため、ロールアウトにもともと必要な複数段階の計算を選択的な記憶伝播に利用する、純粋にRLに基づく構築型の枠組みMemory-in-the-Loop(MiLoop)を提案する。各ロールアウトは学習に解の品質のフィードバックを与えると同時に、過去の表現を伝播する。これにより、外部の解ラベルや学習時の探索空間の枝刈りを使わずに、浅い方策が有効な動的埋め込みを学習できる。具体的には、MiLoopは注意層の前で現在の埋め込みと過去の記憶を融合し、その後に適応的なゲート付き更新を適用する。更新された表現は、現在の意思決定と段階的な再利用の両方を支える。4つのCOPにわたる広範な実験により、100から1000万ノードまでの問題例でMiLoopが一貫して高品質の解を生成することを示し、強い汎化能力を明らかにした。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Constructive neural combinatorial optimization (NCO) has emerged as a promising paradigm that learns to construct solutions to combinatorial optimization problems (COPs) step by step, which reduces reliance on handcrafted rules and enables fast inference. While many methods with dynamic embeddings generalize well, they typically rebuild subproblem representations from scratch at each step using deep attention stacks. Many high-performing methods in this category rely on solution labels or pseudo-labels for efficient training, or on aggressive search space pruning during reinforcement learning (RL). To address these limitations, we propose Memory-in-the-Loop (MiLoop), a purely RL-based constructive framework that leverages the multi-step computation already required by a rollout for selective memory propagation. Each rollout provides solution-quality feedback for learning while propagating historical representations, thereby enabling a shallow policy to learn effective dynamic embeddings without external solution labels or training-time search-space pruning. Specifically, MiLoop fuses current embeddings with historical memory before the attention layers and applies adaptive gated updates afterward. The updated representations support both current decisions and stepwise reuse. Extensive experiments across four COPs demonstrate that MiLoop consistently produces high-quality solutions on instances ranging from 100 to 10 million nodes, highlighting its strong generalization ability.

arXiv ID: 2610.01685 / 要約の誤りについて