質問に応じたモデル重みの更新を分布として予測する
Learning to Predict Distributions over Weight Updates for Test-Time Adaptation
この論文をやさしく読む
ひとことで言うと
質問ごとにモデルの重みをどう変えるかを一つに決めず、複数の候補を生む分布として学習する方法です。
何に役立つ?
追加の実例を渡さず、入力質問からLLMを適応させる仕組みの設計に役立ちます。追加計算を出力文の試行回数ではなく重み更新の候補数へ使う選択肢を示しています。
この研究の面白いところ
同じモデルから答えを何度も生成する方法に対し、少しずつ異なる適応済みモデルを生成して試します。更新が別の質問にも転用できるという観察も報告しています。
どこまで分かった?
要旨には評価課題名、具体的な性能差、計算費用の数値がありません。質問をまたいだ転用は再利用可能な構造を示唆する結果であり、任意の質問への一般化の保証ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ハイパーネットワークは近年、課題の説明や追加の実例などの信号に基づき、実行時に大規模言語モデル(LLM)のパラメーターを動的に適応させることに成功している。本研究では、LLMへの入力質問だけを用いて、どれほどの適応信号を得られるかを問う。これに答えるため、質問を条件とするLoRA推定用のハイパーネットワークを調べる。さらに、パラメーター適応器の点推定だけでなく、取り得るLoRAの分布を生成できる、分布型ハイパーネットワークを導入する。そのために、微分可能なモンテカルロ近似を用いる単純なエンドツーエンド損失を提案し、回帰型と凸結合型を含む複数の分布のパラメーター化を検討する。 結果によると、学習した分布の平均を使うだけでも、決定論的なハイパーネットワークを上回り得る。特に重要なのは、学習した分布が、異なる形のテスト時スケーリングを可能にすることである。固定されたモデルからより多くのトークン列を標本化することだけに追加計算を使う代わりに、重みの更新を標本化し、同じ質問に対して複数の適応済みモデルを得る。検討する重みの標本数を増やすと性能が改善し、対応するトークン標本化による適応の比較手法より高い性能を維持する。最後に、生成した更新は質問をまたいで転用できることが分かり、ハイパーネットワークが、モデルをどう適応させるべきかについて再利用可能な構造を学んでいることが示唆される。これらの結果は、質問を条件とした重み更新の分布が、適応とテスト時スケーリングの両方を支え得ることを示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.
arXiv ID: 2610.01934 / 要約の誤りについて