小さな工程から学ぶカリキュラムでジョブショップ計画を改善する
Curriculum Learning with GNN-based Reinforcement Learning for Job Shop Scheduling
この論文をやさしく読む
ひとことで言うと
小さな工程計画から段階的に学ぶと、大きなジョブショップ計画の学習時間と最適性ギャップが減るかを比較した。
何に役立つ?
グラフニューラルネットワークによる工程スケジューリング方策の学習手順を設計する際に役立つ。実工場での導入効果ではなく、未見の問題事例での評価である。
この研究の面白いところ
目標規模への性能だけでなく、8×8から30×30までの複数規模への汎化と実際の学習時間も比べている。
どこまで分かった?
評価の目標規模は20×20、25×25、30×30である。30×30では平均ギャップが約8.1・8.6ポイント改善し、約50時間短縮したが、他の問題設定への一般化は要旨からは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ジョブショップスケジューリングは難しい組合せ最適化問題であり、グラフニューラルネットワークを使う近年の強化学習は、問題の事例から直接スケジュール方策を学ぶ方法として期待されている。しかし、大きな事例での学習は計算費用が高く、事例の規模をまたいだ汎化も難しい。本論文では、20×20、25×25、30×30という3つの目標規模について、グラフニューラルネットワークによる強化学習にカリキュラム学習を使い、単一規模での学習と比較する。カリキュラム学習では、まず小さな事例で方策を学び、その後より大きな目標規模に段階的に適応させ、前段階で学んだスケジューリングの振る舞いを大きな事例の学習に生かす。モデルは8×8から30×30までの未見の事例で最適性ギャップを使って評価し、全評価規模を通じた汎化と、目標規模への特化の両方を考慮した。結果として、カリキュラム学習は実時間で測った学習時間を一貫して短縮し、目標規模が大きいほど利点が増した。最大の利点は30×30で見られ、全評価規模での平均最適性ギャップを約8.1パーセントポイント、目標規模での平均最適性ギャップを約8.6パーセントポイント減らし、学習時間を約50時間節約した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The job shop scheduling problem is a challenging combinatorial optimization problem, and recent reinforcement learning approaches using graph neural networks have shown promise for learning scheduling policies directly from problem instances. However, training on large instances remains computationally expensive, and generalization across instance sizes remains challenging. This paper studies curriculum learning for graph neural network-based reinforcement learning in the job shop scheduling problem by comparing it with single-size training across three target sizes: 20 x 20, 25 x 25, and 30 x 30. In the curriculum setting, the policy is first trained on smaller instances and then progressively adapted to larger target sizes, allowing scheduling behavior learned in earlier stages to support learning on larger instances. Models are evaluated on unseen instances from 8 x 8 to 30 x 30 using the optimality gap, considering both generalization across all evaluation sizes and specialization on the target size. Results show that curriculum learning consistently reduces wall-clock training time, with larger benefits as the target size increases. The strongest advantage is observed at 30 x 30, where curriculum learning reduces the mean optimality gap across all evaluation sizes by approximately 8.1 percentage points, reduces the target-size mean optimality gap by approximately 8.6 percentage points, and saves approximately 50 hours of training time.
著者のコメント
This paper has been accepted for presentation at the IEEE 10th International Conference on Computational Systems and Information Technology for Sustainable Solutions (CSITSS 2026)
arXiv ID: 2609.28085 / 要約の誤りについて