arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

まれな値を予測する回帰モデルの評価にある見落とし

Exposing Blind Spots in Deep Imbalanced Regression Evaluation

Noah C. Puetz, Jens U. Brandt, Marc Hilbert, Elena Raponi, Thomas Bäck, Thomas Bartz-Beielstein

この論文をやさしく読む

ひとことで言うと

出現頻度の低い値を予測する回帰モデルについて、評価データ、指標、乱数シードの影響を調べ直す。

何に役立つ?

珍しい値での予測精度が重要な回帰システムを評価し、全体平均に隠れる失敗を発見するのに役立つ。

この研究の面白いところ

画像中心の従来評価を時系列の仮想センシングへ広げ、尺度をそろえた指標と複数シードで裾の性能を調べる。

どこまで分かった?

要旨での対象はMuViSの六分野九課題と代表的な六手法。新指標による判断があらゆる分野に妥当だと示したわけではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

深層の不均衡回帰(DIR)は、予測対象の値の分布が大きく偏り、全範囲で信頼できる予測が必要でも、データが多い範囲でモデルの性能が良くなりがちな問題に取り組む。手法は急速に進んだが、評価には三つの盲点がある。画像のベンチマークが中心であること、標準的な多数・中間・少数事例の区分は診断には使えるが意思決定に十分ではないこと、そして分布の裾の安定性を乱数シード間で体系的に調べていないことだ。本研究はこの三点からDIRの評価を見直す。 第一に、六つの物理分野にまたがる九つの時系列外在的回帰課題を含む、マルチモーダルな仮想センシングのベンチマークMuViSでDIRを評価し、対象領域を広げる。ここでは、まれな対象値が運用上意味のある状態に対応することが多い。第二に、バランス調整した平均絶対誤差bMAEを採用し、手法やデータセットをまたいだ意思決定に使える、尺度で正規化した指標bMASEを導入する。第三に、代表的な六つのDIR手法を複数の乱数シードで繰り返し評価し、DIRが重視する分布の裾がシードによる変動に特に敏感だと示す。 結果として、通常の仮想センシングモデルは全体のMAEでは隠れる大きな裾の性能低下を示す。既存のDIR手法はバランスを考慮した性能を改善できる一方、マルチモーダル時系列への転用効果は一様ではない。また、裾の領域の不安定さは現在のDIR評価では見落とされがちである。著者らは、これらの知見と公開コードが、まれな値も一般的な値と同じように確実に扱う回帰システムの研究に再現可能な基盤を与えると述べる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Deep Imbalanced Regression (DIR) addresses a common failure mode of regression models: target distributions are highly non-uniform, causing models to perform best in densely populated target regions even when reliable performance is required across the full target range. Despite rapid methodological progress, DIR evaluation remains constrained by three blind spots: it is dominated by image-based benchmarks, its standard many-/medium-/few-shot protocol is diagnostic but not decision-complete, and tail-region stability across random seeds has not been systematically evaluated. We revisit DIR evaluation along these three axes. First, we broaden the data domain by evaluating DIR on a multimodal virtual sensing benchmark (\textsc{MuViS}) with nine time-series extrinsic regression tasks across six physical domains, where rare target values often correspond to operationally meaningful regimes. Second, we adopt balanced MAE (\emph{bMAE}) and introduce balanced Mean Absolute Scaled Error (\emph{bMASE}), a scale-normalized metric for decision-complete comparison across methods and datasets. Third, through a repeated reevaluation of six representative DIR methods across multiple random seeds, we show that the tail regions targeted by DIR exhibit particularly high sensitivity to seed-level variability. Our results show that standard virtual-sensing models exhibit substantial tail degradation hidden by global MAE, that existing DIR methods can improve balanced performance but transfer unevenly to multimodal time-series data, and that tail-region instability remains a largely hidden failure mode under current DIR evaluation practice. Together, these findings and our publicly available code provide a reproducible basis for future DIR research toward regression systems that capture rare target regimes as reliably as common ones.

arXiv ID: 2609.25152 / 要約の誤りについて