タスク間の共通構造が文脈内学習の必要データ量を減らす仕組み
Transformers as Cross-Task Learners: Shared Structure Drives Sample Efficiency in In-Context Learning
この論文をやさしく読む
ひとことで言うと
関連する複数のタスクに共通の低次元構造があると、Transformerが短い例から未知のタスクを学びやすくなる理由を数学的に調べています。
何に役立つ?
文脈内学習に必要な事前学習タスク数やプロンプト長を考えるための理論的な指針になります。実システムでの性能改善を実証したという主張ではありません。
この研究の面白いところ
タスクを特定のパラメータ式で表さず、被覆数から基準関数を作り、その手順をSoftmax注意機構のTransformerで近似する構成まで示しています。
どこまで分かった?
要旨が述べる成果は構成と誤差上界を中心とする理論解析です。実データや実運用のモデルでの検証については要旨に記載がありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Transformerは事前学習中に幅広いタスク群をまとめて学び、短いプロンプトだけで未知のタスクに適応することで高い性能を示す。しかし、この現象の厳密な数学的・統計的理解はまだ限られている。本研究は、Transformerがタスク間に共通する構造をどう利用し、その構造が文脈内学習(ICL)のサンプル複雑性にどう影響するかを調べる。 具体的には、指定した距離の下での被覆数を用いてタスク空間の複雑さを特徴付け、明示的なパラメータ表現を必要とせずに、タスク間に共通する低次元構造を定量化する。得られた被覆から基準となる関数群を用意し、タスクの同定と評価を行う手順を導入する。文脈中の観測によって未知のタスクを基準関数群の中で絞り込み、該当する基準関数の問い合わせ点での評価値を集約して応答を予測する。 近似については、この手順を近似するSoftmax注意機構付きのTransformerを明示的に構成する。汎化については、事前学習タスクの数とプロンプトの長さの影響を分けた誤差上界を導く。事前学習タスク数に関するスケーリングはタスク空間と入力領域の内在次元に支配され、十分な数のタスクが得られると、プロンプト内の文脈の長さへの依存は次元に依存しなくなる。著者らの知る限り、一般的な非線形タスク群のタスク間複雑さを定量化し、その低次元構造を利用してICLを行うTransformerを明示的に構成した初めての研究である。この理論は、関連するタスクを共同で事前学習すると文脈内での汎化が改善する仕組みを定量的に説明する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Transformers achieve remarkable performance by jointly learning broad families of tasks during pretraining and adapting to unseen tasks from only a short prompt. Yet a rigorous mathematical and statistical understanding of this phenomenon remains limited. This paper aims to study how Transformers exploit shared cross-task structure and how this structure affects the sample complexity of in-context learning (ICL). Specifically, we characterize task-space complexity through covering numbers under a prescribed metric, thereby quantifying the low-dimensional cross-task structure without requiring an explicit parametric representation. The resulting cover provides a set of anchor functions, which we use to introduce a task-identification-and-evaluation procedure: context observations localize an unseen task among the anchor functions, and the response at a query is predicted by aggregating the corresponding anchor function query evaluations. For approximation, we explicitly construct a Transformer with Softmax attention to approximate this procedure. For generalization, we derive an error bound that separates the effects of the number of pretraining tasks and the prompt length. The scaling with respect to the number of pretraining tasks is governed by the intrinsic dimensions of the task space and input domain; once sufficiently many tasks are available, the dependence on the prompt context length becomes dimension-free. To the best of our knowledge, this is the first work to quantify cross-task complexity for general nonlinear task families and explicitly construct a Transformer that exploits their low-dimensional structure to perform ICL. Our theory provides a quantitative explanation of how joint pretraining across related tasks improves in-context generalization.
arXiv ID: 2609.29060 / 要約の誤りについて