不均衡な表データの学習行数を減らす抽出法
CRISP: Scalable Importance-Stratified Coresets for Imbalanced Tabular Learning
この論文をやさしく読む
ひとことで言うと
不正検知などの不均衡な表データから、学習に使う行を大幅に減らす方法を示した。
何に役立つ?
勾配ブースティング木を繰り返し学習する際の計算量削減に役立つ可能性がある。
この研究の面白いところ
本番データで総行数を93.2%減らしながら、全データの平均適合率の99.7%を保った。
どこまで分かった?
Sparkovの低い削減率では結果が混在したため、すべてのデータと削減率で最良とは言えない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模でクラスが不均衡な表データでは、勾配ブースティング木の繰り返し学習に費用がかかる。既存のコアセット法では、多数派の例をほとんど除くと精度を失いやすい。本研究はCRISPを提示する。代理モデルのスコアを分位の層に分け、負例の採用予算を層ごとに割り当てる線形時間の方法である。標本重みで採用確率の違いを補正する。 本番の不正検知データで負例を95%減らしたとき、2,500万行のうち約170万行で学習し、全データ学習時の平均適合率の99.7%を保持した。学習行数全体では93.2%の削減である。公開データCriteoPrivateAdsでは、多数派を90%から99.4%減らす各試験率で、平均適合率の平均が最も高かった。Sparkovでは低い削減率で結果が混在したが、99.2%と99.4%では最も高い平均を得た。要素を除いた比較から、予算配分と採用確率の逆数による重み付けが、本番データでの改善の主因と分かった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large imbalanced tabular datasets make repeated gradient-boosted tree training expensive. Existing coreset methods often lose accuracy when most majority examples are removed. We present CRISP (Coreset Reduction via Importance-Stratified Pruning), a linear-time method that allocates a negative-class budget across quantile strata of a proxy-model score. Sample weights account for unequal inclusion probabilities. At 95% negative-class reduction on a production fraud dataset, CRISP trains on approximately 1.70M of 25M rows and retains 99.7% of full-data Average Precision. This is a 93.2% reduction in total training rows. On public CriteoPrivateAds, CRISP has the highest mean Average Precision at each tested rate from 90% to 99.4% majority reduction. Sparkov results are mixed at lower rates, but CRISP has the highest mean at 99.2% and 99.4%. Ablations identify budget allocation and inverse-propensity weighting as the main sources of the production-dataset gain.
arXiv ID: 2609.26962 / 要約の誤りについて