代替商品の在庫を考慮した動的価格設定を方策学習で近似
Scalable Dynamic Pricing of Substitutable Products through Structure-Guided Policy Learning
この論文をやさしく読む
ひとことで言うと
在庫が限られ、客が別の商品へ乗り換える状況で、複数商品の価格を決める方法です。厳密な計算の代わりに、在庫から価格や在庫の価値を予測する方策を学びます。
何に役立つ?
考えられる用途は、商品数や在庫状態が多く、厳密な動的計画が難しい収益管理です。航空業界を想定した大規模な数値問題でも比較しています。
この研究の面白いところ
価格を直接学ぶ方法と、在庫の機会費用を学んで既知の価格則へ渡す方法を比較しています。需要モデルがよく当てはまる場合と、不均質な場合で適した構成が変わります。
どこまで分かった?
0.4%未満は最適解を計算できる小規模問題での平均ギャップです。大規模問題は比較ベンチマークに対する改善であり、同じ最適性ギャップが保証されたわけではありません。実店舗や実便での導入実験は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
問題設定:商品ごとに有限の在庫を持つ代替可能な商品の動的価格設定を研究する。顧客の代替選択によって商品間の価格判断が結び付き、在庫状態があるため、現実的な規模では厳密な動的計画法が扱いきれなくなる。 方法と結果:動的計画法を、在庫状態から価格判断への統計的な写像に置き換える、多項ロジット(MNL)に導かれた2つの方策学習法を開発する。第1の方法は価格を直接学び、第2の方法は在庫の機会費用を学んで、最適なMNL価格設定則を使い価格へ変換する。両方の方策を意思決定重視の学習で訓練する。また、方策を学習するために、将来を見越した顧客選択の教師目標を効率よく生成する手法を提案する。 経営上の示唆:提案手法を広範に数値評価する。最適な動的計画法を計算できる小規模な問題では、学習した方策の平均最適性ギャップは0.4%未満となる。航空業界を念頭に置いた大規模な問題では、比較した収益管理のベンチマークを一貫して上回る。2つの構成の比較は、モデル構造の役割も明らかにする。需要がMNLでよく記述される場合には、MNLの価格設定の特徴付けを利用する方法が特に有効である。一方、不均質な混合MNL需要の下では、価格を直接学ぶ方法の方が柔軟性に優れる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Problem definition: We study dynamic pricing of substitutable products with finite, product-specific inventories. Customer substitution couples pricing decisions across products, while the inventory state makes exact dynamic programming intractable at realistic scale. Methodology / results: We develop two MNL-guided policy-learning approaches that replace the dynamic program with a statistical mapping from inventory states to pricing decisions. The first learns prices directly, while the second learns inventory opportunity costs and converts them into prices using the optimal MNL pricing rule. Both policies are trained using decision-focused learning. We also propose an efficient method to generate anticipative customer-choice targets to train our policies. Managerial implications: We conduct an extensive numerical evaluation of our approaches. On small instances for which the optimal dynamic program can be computed, the learned policies achieve average optimality gaps below 0.4%. On larger airline-motivated instances, they consistently improve on the tested revenue-management benchmarks. The comparison between the two architectures also highlights the role of model structure: using the MNL pricing characterization is particularly effective when demand is well described by MNL, while directly learning prices provides greater flexibility under heterogeneous mixed-MNL demand.
arXiv ID: 2609.24605 / 要約の誤りについて