arXiv論文メモ
新着一覧
cs.LG / cond-mat.dis-nn / stat.ML · 査読状況未確認

非線形課題の文脈内学習で2種類の注意モデルを比較する

In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners

Haotian Gu, Yizhou Xu, and Lenka Zdeborová

この論文をやさしく読む

ひとことで言うと

実例を文章内などに示して課題を解かせる学習について、非線形変換を注意処理の前に置くか後に置くかで、性能がどう変わるかを理論と数値実験で比較しています。

何に役立つ?

データ量やタスクの多様性、提示できる実例数に応じて、モデルの構造の向き不向きを考える基礎になります。

この研究の面白いところ

一方が常に優れるとするのではなく、複数の条件を変えた相図として有利な領域を整理します。学習時と推論時の文脈長の影響も区別しています。

どこまで分かった?

対象は単一指標型の非線形タスクと2種類の1層アテンションモデルです。レプリカ法による予測を数値実験で確かめていますが、一般の大規模言語モデルで同じ相図になることを実証した要旨ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文脈内学習(ICL)では、事前学習済みモデルがパラメータを更新せずに、提示された実例からタスクを推定できる。既存理論の多くは線形の目標関数に焦点を当てているが、本論文では、同じ単一指標型タスク群に対して2種類の1層アテンション構造を比較し、非線形の場合を調べる。カーネル学習器は、まず固定された非線形特徴写像で入力を変換してから線形アテンションを適用する。一方、特徴学習器は元の入力にアテンションを適用した後、学習された非線形読み出しを行う。 レプリカ法を用い、事前学習データ量、タスク集合の多様性、学習時と推論時の文脈長の効果を残したまま、記憶誤差と汎化誤差を予測する式を導く。得られた予測は、広い範囲の条件で数値実験によく一致する。解析からは、事前学習データ量、タスクの多様性、文脈長が変わるとき、どちらの構造が有利になるかを特徴付ける相図が得られる。さらに、2つの学習器では文脈長に対するスケーリングが質的に異なることを特定する。これらの結果は、構造の選択がデータセットとどう相互作用し、非線形の文脈内学習を左右するかを明らかにする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tasks. A kernel learner first maps inputs through a fixed nonlinear feature map and then applies linear attention, whereas a feature learner applies attention to the original input, followed by a learned nonlinear readout. We derive predictions for their memorization and generalization errors using the replica method, retaining the effects of pretraining size, task-pool diversity, and training and inference context lengths. The resulting predictions closely match numerical experiments across a broad range of regimes. Our analysis yields phase diagrams that characterize when each architecture is advantageous as the amount of pretraining data, task diversity, and context lengths vary. We further identify qualitatively different context-length scalings for the two learners. Together, these results clarify how architectural choices interact with the dataset and govern nonlinear in-context learning.

arXiv ID: 2610.01712 / 要約の誤りについて