arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

学習済みモデルの層を再利用して回答候補を多様化

Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models

Akshit Singh, Shyam Marjit, Wei Lin, Leonid Karlinsky, M. Jehanzeb Mirza

この論文をやさしく読む

ひとことで言うと

同じモデルに回答させる際、一部の層を繰り返し通すことで計算経路を変え、正しい回答を含む候補群を得やすくする方法です。

何に役立つ?

複数回答から正答を得る処理や、モデルが自分で生成した回答を使うテスト時学習に役立ちます。後者でも精度改善を報告しています。

この研究の面白いところ

出力時の乱数だけでなく、モデル内部の計算経路を変えて多様性を作ります。貪欲復号でも候補の網羅性が改善した点が特徴です。

どこまで分かった?

平均6.58ポイントの改善はpass@9で、9候補の少なくとも1つが正しいかをみる結果です。単一回答の精度改善と同じ意味ではなく、候補数が同じでも計算量が等しいとは要旨に記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

テスト時の計算量拡大では、重みを固定したモデルから複数の応答をサンプリングし、よりよい回答を得ることがしばしば目指される。しかし従来の温度サンプリングでは、全ての候補が同じ固定された計算経路を通って生成される。本研究では、デコーダ層の一部のブロックを再利用し、異なる順伝播計算を通じて候補を生成する、追加学習不要の手法「アーキテクチャサンプリング」を導入する。ブロックの位置と反復回数を変えることで、モデルの重みを更新したり補助パラメータを加えたりせずに、計算の多様性を導入する。 Qwenの5つのチェックポイントと12のマルチモーダルベンチマークにわたり、候補数を同じ9個にそろえた条件で、標準経路の温度サンプリングに比べてpass@9が平均6.58ポイント向上した。初期の層を再利用した場合に最も大きな改善が得られ、候補が正答を含む範囲の改善は、貪欲復号でも維持された。生成された候補は語彙の重複が少なく、ラベルを使わないテスト時強化学習のロールアウトとして用いると精度が向上する。これらの知見は、本手法の利点が候補の網羅性にとどまらず、モデル自身の出力からの学習をより効果的にすることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.

arXiv ID: 2610.01687 / 要約の誤りについて