選択モデルの予測の違いを確率的混同行列で見る
Using machine learning metrics to provide deeper insights into the performance of choice models
この論文をやさしく読む
ひとことで言うと
選択モデルの良さを総合的な当てはまりだけでなく、どの選択肢をどれと混同するかでも評価します。
何に役立つ?
同じ程度の対数尤度を持つモデルの違いを見つけ、仕様の見直しや予測性能の検討に役立ちます。
この研究の面白いところ
観測された選択肢ごとの予測確率を平均した確率的な混同行列を用い、機械学習の評価指標を選択モデルへ移します。
どこまで分かった?
評価の見方を増やす提案です。どのモデルが常に最良かを決める単一指標ではなく、選択肢間のトレードオフを明らかにします。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
選択モデリング(CM)の分野では、機械学習(ML)手法への関心が高まっている。これまでの焦点は、対照的な両アプローチの性能比較や、機械学習手法から行動に関する洞察を得やすくすることにあり、一方の分野の発想を他方へ移すことにはあまり向けられていなかった。本論文では、モデルの性能評価におけるMLからCMへの知識移転に焦点を当てる。 CMでは通常、対数尤度や関連指標でモデル性能を評価するが、これらは全体的な適合度に重点を置く集約的な指標である。一方、MLでは選択肢ごとの誤分類と正しい分類に重点を置き、結果をより細かく捉える。両者を結びつけるため、確率版の混同行列の利用を検討する。この行列は、実際にどの選択肢が選ばれたかを条件として、モデルが各選択肢を予測する確率を全選択課題にわたって平均した値を示す。これにより、古典的な選択モデルとMLアルゴリズムの双方について、確率的なML指標を計算できる。 全体的な適合度と選択肢ごとの予測を併せて性能を分析した結果、対数尤度が似たモデルでも混同行列は大きく異なり、集約指標では捉えられない確率パターンの違いが明らかになった。この枠組みは、モデルがどの選択肢を系統的に混同するかを特定し、選択肢間のトレードオフを明らかにすることで、モデルの定式化を導く可能性がある。さらに、標本外でこれらの行列と指標を評価すると、予測性能に大きく影響する、選択肢ごとの予測の変化が明らかになる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model performance is typically evaluated using log-likelihood and related indicators, which are aggregate fit metrics that focus on overall fit. Conversely, in ML, the focus is on alternative-level misclassifications and correct classifications, which provide a more nuanced view of the results. To bridge these approaches, we explore the use of a probabilistic version of the confusion matrix, which reports the average probability of the model predicting each alternative, conditional on which alternative was observed to be chosen, across all choice tasks. This enables the computation of probabilistic ML metrics for both classic choice models and ML algorithms. We analyse model performance jointly in terms of overall fit and alternative-level predictions. Our findings demonstrate that models with similar log-likelihood can exhibit substantially different confusion matrices, revealing different probability patterns that aggregate metrics cannot capture. This framework identifies where models systematically `confuse' alternatives, highlighting trade-offs between alternatives, and potentially guiding model specification. Furthermore, evaluating these matrices and metrics out-of-sample reveals alternative-level prediction shifts that significantly impact forecasting performance.
arXiv ID: 2609.20655 / 要約の誤りについて