arXiv論文メモ
新着一覧
stat.ML / cs.LG · 査読状況未確認

疎なガウス過程で使う基底関数の選び方

On Basis Function Selection for Sparse Gaussian Process Regression

Marnix Van Soom, Ivan De Boi

この論文をやさしく読む

ひとことで言うと

計算量を抑えたガウス過程モデルで、先頭の基底関数を機械的に使う代わりに、重要な関数を選ぶ方法を比較しました。

何に役立つ?

限られた計算予算でガウス過程回帰の精度を上げるための基底選択に役立ちます。

この研究の面白いところ

情報理論から三つの選択基準を導き、6件の回帰ベンチマークと3種類の基底群で評価しています。

どこまで分かった?

要旨の性能評価は指定された6件のベンチマークと3種類の基底群に基づきます。すべてのデータやカーネルでの改善は示していません。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

疎なガウス過程は、入力空間上の固定基底関数{φ_j}による適切な展開でカーネルを置き換え、O(N)の推論を可能にする。計算予算がM≪Nのとき、通常は基底の最初のM個を残して切り詰める。しかし、この形式では、手元のデータに重要なM個だけを選ぶことも妨げられていない。そうすれば信号のない基底に予算を費やさずに済むが、候補の順位付けの基準が必要になる。 本研究は、基底関数の選択を情報理論の観点から捉え、三つの基準を提案する。それぞれ選択時にデータがない場合、事前情報がない場合、その中間の場合に対応する。次に、Hilbert空間ガウス過程(HSGP)、変分フーリエ特徴(VFF)、変分誘導球面調和関数(VISH)の三つの基底群を使い、UCIの回帰ベンチマーク6件で切り詰めと選択方法の性能を比較する。データを使わない基準は安全な標準選択肢であり、HSGP、VFF、VISHのすべてで切り詰めと同等またはより良い結果を示した。VISHでは大幅な改善があり、この基底群向けに最近開発された選択ヒューリスティックも上回った。データを考慮する「事前情報なし」と「中間」の基準は、三群のうち実務で最も広く使われるHSGPで特に、切り詰めより大幅に良い結果を示した。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Sparse Gaussian processes achieve $O(N)$ inference by replacing the kernel with an appropriate expansion in a fixed basis $\{\phi_j\}$ on the input space. Given a compute budget $M \ll N$, practitioners conventionally truncate the basis to its first $M$ entries. Nothing in the formalism, however, prevents one from selecting only those $M$ basis functions that matter for the data at hand. This would avoid spending budget on basis functions where there is no signal, but it requires a criterion for ranking the candidates. We propose three such criteria derived from an information-theoretic view of the basis-function selection problem. Each criterion matches a different state of knowledge at selection time: a no-data state, a no-prior state, and an in-between state. We then study the performance of truncation versus selection strategies on six UCI regression benchmarks across three basis families: Hilbert-space Gaussian processes (HSGP), variational Fourier features (VFF), and variational inducing spherical harmonics (VISH). We observe that the no-data criterion is a safe default, matching or improving on truncation for HSGP, VFF and VISH, with substantial gains for VISH and improvements over a recently developed selection heuristic for that basis family. The data-aware no-prior and in-between criteria provide substantial gains over truncation specifically for HSGP, which is the most broadly used of the three families in practice.

著者のコメント

18 pages, 8 figures

arXiv ID: 2609.26624 / 要約の誤りについて