観測ごとに重要な行を見つける注意機構付き回帰
Diagonalized Attention for Individualized Regression: Latent-Row Localization and Prediction
この論文をやさしく読む
ひとことで言うと
文章ごとに重要な単語が違うようなデータに対し、共通の場所だけを見るのではなく、各観測の重要部分を選んで予測します。
何に役立つ?
局所的な特徴を並べた行列データで、予測と重要箇所の特定を同時に行うための枠組みになります。実データでは感情分類への適用が報告されています。
この研究の面白いところ
選ぶ行は観測ごとに変わる一方、回帰効果は共有します。新しい観測の正解を知らなくても重要行を探せる構造と、注意機構の統計的な説明を結び付けています。
どこまで分かった?
理論はスコアの分離と集中などの条件を課した存在定理です。任意のデータや学習手順での成功を保証するものではなく、要旨には実データの精度改善幅は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
現代のテキストや画像の表現は、行がトークン、画像パッチ、その他の局所特徴ベクトルに対応する行列の形を取ることが多い。予測に役立つ情報はしばしば疎でありながら標本ごとに異なるため、非ゼロの位置を共通に仮定する従来のスパース回帰法は、この異質性に適していない。 本論文では、行列値の共変量に対する個別化スパース回帰の枠組みを定式化する。各観測はそれぞれ固有の重要な行を持つ一方、対応する回帰効果は母集団全体で共有される。このモデルを推定するため、クエリとキーのスコアによって標本固有の信号行を特定し、バリュー行列を後段の回帰に用いる、対角化された注意機構を導入する。提案法のパラメータ次元は標本数に依存せず、新たな観測についても目的変数を知らずに重要な行を特定できる。 適切なスコア分離条件と集中条件の下で、単一ヘッドおよび複数ヘッドの対角化注意モデルが潜在的な行を高確率で復元し、予測リスクの上界を与えることを示す存在定理を確立する。これにより本理論は、注意に基づくスコア付けが、異質な行列値データの中の標本固有の信号をどのように特定するかを統計的に説明する。 シミュレーションでは、標本数、次元、信号の個数を変化させ、回帰およびモデルの誤指定を伴う分類において、高い予測性能と特定性能を示した。実データの感情分析では、分類精度の改善と解釈可能なトークン選択が示された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Modern text and image representations are often matrix-valued, with rows corresponding to tokens, patches, or other local feature vectors. Predictive information is often sparse but sample-specific, making classical sparse regression methods with a common support poorly suited to this heterogeneity. This paper formalizes an individualized sparse regression framework for matrix-valued covariates in which each observation has its own rows of interest, while the associated regression effects are shared across the population. To estimate this model, we introduce a diagonalized attention mechanism that uses query--key scores to localize sample-specific signal rows and a value matrix for downstream regression. The proposed method has a parameter dimension independent of sample size and can identify rows of interest for new observations without their responses. We establish existence theorems showing that, under suitable score-separation and concentration conditions, single-head and multi-head diagonalized attention models recover the latent rows with high probability, yielding prediction risk bounds. Our theory therefore provides a statistical explanation of how attention-based scoring localizes sample-specific signals in heterogeneous matrix-valued data. Simulations demonstrate strong prediction and localization in regression and misspecified classification across varying sample sizes, dimensions, and signal cardinalities. Real sentiment analyses show improved classification accuracy and interpretable token selection.
arXiv ID: 2609.21320 / 要約の誤りについて