継続学習で専用部品を置く層を脳活動モデルから選ぶ
Two Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize
この論文をやさしく読む
ひとことで言うと
新しいタスクに合わせる部品を、画像モデルのどの深さに置くかを選ぶ研究です。脳活動を予測する固定モデルを使った選択が、総当たり探索で有利だった中間層と重なりました。
何に役立つ?
タスクが増えるたびに必要になるアダプターの保存容量を抑える判断に役立ちます。Split ImageNet-Rでは容量60%で最終精度の差を1.5ポイント以内に抑えた結果が示されています。
この研究の面白いところ
安価な重み・活性化の指標が深い層を選ぶのに対し、脳活動モデルを介した指標は探索で有利な中間層を選びます。学習器外部の表現を配置の判断に使う点が特徴です。
どこまで分かった?
評価は3種類のViT-B/16で、配置探索との直接比較はAugRegとiBOTの2種類です。要旨は、この結果が任意のモデルに成り立つとは示しておらず、人間の脳と同じ仕組みで学習することを実証したものでもありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
事前学習済み視覚Transformerのすべてのブロックにタスク専用アダプターを保持する継続学習器では、タスク数に比例して保存容量が増える。専用アダプターを少数のブロックだけに置けば増加を抑えられるが、どこに配置すべきかが問題となる。本研究では、この問題を二つの観点から調べる。アルゴリズムの観点では、連続する4ブロックの全配置を学習すると、最終精度は逆U字形の関係を示す。中間の深さで最大となり、配置によって最大3.5パーセントポイント変わる一方、重みのスペクトルや活性化統計に基づく低コストの基準は最も深いブロックを選ぶ。 神経科学の観点では、視覚皮質の階層構造と中間段階の可塑性を手掛かりに、学習器の外部で得た測定によって、配置探索を行わずに層の専門化を導けるかを問う。LS-Bは、人間の12の視覚領域を表す固定済みfMRI符号化モデルを通じて最初の数タスクを観察し、安定した構造に対して読み出しがタスク間で最も変化するブロックに、タスク専用容量を一度だけ割り当てる。3種類のViT-B/16バックボーンで、LS-Bは安定した、バックボーン固有の配置を得る。配置探索を行ったAugRegとiBOTの二つでは、選択されたブロックが探索で見つかった中間深度の領域と重なる。保存容量と観察予算を揃えると、選択されたブロックは最も浅い4ブロックと最も深い4ブロックの構成を上回る。Split ImageNet-Rでは、LS-Bは全ブロック版BiLoRAのアダプター保存容量の60%を使いながら、最終精度の差を1.5パーセントポイント以内に抑える。割り当てにはラベルも誤差逆伝播も不要で、実行時間の増加は0.6%未満であり、バックボーン固有の皮質領域のパターンが現れる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Continual learners that keep a task-specific adapter in every block of a pre-trained vision transformer accumulate storage linearly with the number of tasks; keeping task-specific adapters in only a few blocks curbs this growth but raises the question of where to place them. We investigate this question from two perspectives. Algorithmically, training all contiguous four-block placements yields an inverted U: final accuracy peaks at intermediate depth and varies by up to 3.5 percentage points (pp), while inexpensive criteria based on weight spectra or activation statistics favor the deepest blocks. From neuroscience, the hierarchical organization and intermediate-stage plasticity of the visual cortex motivate us to ask whether a measurement taken outside the learner can guide layer specialization without placement search. LS-B observes the first tasks through a frozen fMRI encoding model of twelve human visual areas and commits task-specific capacity once to the blocks whose readouts vary most across tasks relative to their stable structure. Across three ViT-B/16 backbones, LS-B yields stable, backbone-specific allocations. On the two backbones with placement search, AugReg and iBOT, the selected blocks overlap the intermediate-depth region identified by search. Under matched storage and observation budgets, the selected blocks outperform the shallowest and deepest four-block configurations. On Split ImageNet-R, LS-B uses 60% of full-BiLoRA adapter storage while remaining within 1.5 pp of its final accuracy. The allocation requires no labels or backpropagation, adds under 0.6% runtime, and exhibits backbone-specific cortical signatures.
著者のコメント
21 pages, 12 figures
arXiv ID: 2610.01590 / 要約の誤りについて