arXiv論文メモ
新着一覧
stat.ME / stat.CO · 査読状況未確認

行と列の両方向に分割して大規模データの変数を選ぶ

A Bayesian Bi-Directional Splitting Framework for Variable Selection in Large Datasets

Aaron Coats, Vinny Davies, Mayetri Gupta

この論文をやさしく読む

ひとことで言うと

行数も列数も多い表形式データを両方向へ分割し、重要な変数をベイズ的に選ぶ方法です。

何に役立つ?

大量の標本と多数の説明変数が同時にあるとき、計算負担を抑えて信号を見つけるために役立ちます。H3N2インフルエンザのデータにも適用しています。

この研究の面白いところ

各小分けデータを独立に並列解析し、2段階の合意処理で結果を統合します。標本数だけ、変数数だけの分割とは異なり両方の負担へ対応します。

どこまで分かった?

分割は情報の損失を伴います。実験では選択性能を保ちつつ計算上の利得を得ていますが、要旨には速度や選択精度の数値、あらゆる分割での保証はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

現代の表形式データは、サンプル数と共変量数の両方で大規模化している。その結果生じる計算負荷は、ベイズ変数選択に大きな課題をもたらす。観測数が多い場合や共変量空間が高次元の場合にベイズ推論を拡張する研究は豊富にあるが、両方が同時に大きい実用的な設定で、拡張可能なベイズ変数選択を扱う研究は比較的少ない。 本論文は、行数と列数の両方が多いデータを効率的に解析する、新しいベイズ変数選択の枠組みを提示する。分割統治法を用い、データを両方向にバッチ分割して、各バッチを独立に並列解析した後、2段階の合意形成手順で結果を統合する。 実験は、良好な変数選択性能を保ちながら計算上の利得が得られることを示す。データ分割に伴う情報の損失があっても、雑音のある高次元の設定で、関連する信号を正しく特定できた。調整パラメータの選択や効果的なデータ分割戦略についての推奨を含む実用的な指針も示し、続いてH3N2インフルエンザのデータセットへの実応用を提示する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Modern tabular datasets are becoming increasingly large, both in the number of samples and covariates, posing significant challenges for Bayesian variable selection due to the resulting computational burden. While there is extensive literature on scaling Bayesian inference to large numbers of observations or high-dimensional covariate spaces, comparatively little work addresses scalable Bayesian variable selection when both dimensions are large simultaneously in a practical setting. This paper presents a novel Bayesian variable selection framework for efficiently analysing data with a large number of both rows and columns. The proposed framework operates via a divide-and-conquer approach, splitting data into batches along both directions and analysing each batch independently in parallel, after which the results are combined together in a two-phase consensus procedure. Experiments demonstrate the computational gains while retaining strong variable selection performance, successfully identifying relevant signals in noisy, high-dimensional settings despite the loss of information induced by data splitting. Practical guidelines are also provided, including recommendations for tuning parameter choices and effective data partitioning strategies, followed by a real application to an H3N2 influenza dataset.

arXiv ID: 2609.20000 / 要約の誤りについて