arXiv論文メモ
新着一覧
stat.ME / cs.LG / stat.ML · 査読状況未確認

数値や分類が混在する結果の同時分布を学ぶ回帰法

Generalized Engression Models

Xinwei Shen, Zijian Guo, Francis Bach

この論文をやさしく読む

ひとことで言うと

数値、分類、順序などが混ざった複数の結果を、互いの関係も含めて一つの確率分布として学ぶ方法です。

何に役立つ?

複数種の出現状況や、異なる型の健康指標をまとめて予測する統計モデルに使うことが考えられます。個々の結果の平均だけでなく、結果同士の依存も扱える点が目的です。

この研究の面白いところ

データ型に合わせたリンク関数を使いながら、確率的な摂動で不連続な部分も勾配学習できるようにします。単一成分の評価を維持しつつ、同時分布の評価で改善を報告しています。

どこまで分かった?

要旨の応用例は242種の生態データと17次元の健康アウトカムです。具体的な改善量や、あらゆる混合データでの優位性は示されていません。健康データへの適用は、臨床介入の効果を示すものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

共変量が与えられた下で、複数の成分を持つ結果の条件付き分布を推定する問題を考える。各成分は連続値、二値、カテゴリ、順序、順位であり得て、条件付きでも互いに依存する。結果の型ごとにさまざまな統計手法が開発されてきたが、その多くは結果ベクトルの同時分布ではなく、各成分の平均など条件付き分布の要約を対象としている。 本研究では、任意の型の結果を対象とする、統一的なノンパラメトリック分布回帰の枠組み、一般化engressionモデルを開発する。提案手法は、スコアリング規則に基づく深層生成モデルengressionを基盤とし、データ型ごとのリンク関数と、損失をなめらかにする確率的摂動を導入する。これによって、リンクが不連続でも勾配による学習が可能になる。連続、離散、混合の結果に関する普遍的な表現の結果を確立する。 シミュレーションと二つの応用、すなわち群集生態学のベンチマークにおける242種と、17次元の混合型の健康アウトカムにおいて、本手法は周辺分布のスコアでは型別のモデルと同等で、同時分布ではそれらを改善し、専用に作られた最先端の種の同時分布モデルと同等以上の性能を示した。Pythonのソフトウェアを利用できる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We consider estimating the conditional distribution of a multivariate outcome given covariates when its coordinates may be continuous, binary, categorical, ordinal or rankings, and are conditionally dependent on one another. Different statistical methods have been developed for each outcome type, and most of them target a summary of the conditional distribution, such as the mean of each coordinate, rather than the joint distribution of the outcome vector. We develop generalized engression models, a unified nonparametric distributional regression framework for outcomes of any type. The proposed method builds upon engression, a scoring-rule-based deep generative model, and introduces a data-type-specific link function and a stochastic perturbation that smooths the loss, enabling gradient-based training even with discontinuous links. We establish universal representation results for continuous, discrete and mixed outcomes. In simulations and in two applications, 242 species in a community ecology benchmark and a 17-dimensional mixed-type health outcome, the method matches type-specific models on marginal scores, improves on them on the joint distribution, and matches or exceeds purpose-built state-of-the-art joint species distribution models. Software is available in Python.

arXiv ID: 2610.01823 / 要約の誤りについて