arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

モデル構造に合わせた重み摂動でRandOptの候補数を削減

Modular Norm RandOpt: Population-Efficient Ensembling through Architecture-Aware Perturbations

Kirato Yoshihara, Hiroaki Hamade

この論文をやさしく読む

ひとことで言うと

言語モデルの重みを少し変えた候補を組み合わせる際、部品ごとに摂動の大きさを調整する方法です。

何に役立つ?

多数の候補を試す勾配なしの探索で、計算資源を抑える設計の参考になります。要旨では複数の課題とモデル系列で評価しています。

この研究の面白いところ

Countdownで候補を3分の1、GSM8Kで少なくとも12分の1にしながら比較手法を上回りました。ただし改善の説明は単純な候補分布の裾だけでは足りません。

どこまで分かった?

性能差は指定された七課題とモデル群での評価です。すべてのモデルや課題で同じ候補削減率が得られるとは示していません。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

RandOptは重みに摂動を加えた言語モデルをサンプリングし、上位候補の多数決を使うが、全体に共通する摂動の大きさはモジュールごとに異なる幾何学的性質を無視する。本研究は、候補の選別と投票は維持しつつ、モジュールごとの自然なノルムと調整した尺度を用いる、構造を考慮したサンプリング法Modular Norm RandOptを提案する。Countdownでは候補数を3分の1、GSM8Kでは少なくとも12分の1にしてRandOptを上回り、実時間も節約した。七つの課題と0.5B~3Bの三規模のQwenモデルによる評価では、すべての規模でCountdown、GSM8K、MATH-500の平均正答率がRandOptより高かった。CountdownとGSM8Kでは、改善はLlama 3.2 3BとGemma 3 4Bにも及んだ。Qwen2.5-1.5Bでは、主な実行に使う評価予算が同程度の条件で、二つの課題とも反復型の基準手法よりアンサンブルの平均正答率が高かった。GSM8Kの裾の密度に基づく診断から示唆される候補削減は1.2~1.8倍にとどまり、アンサンブル改善の大部分は正答する専門候補をより有利に支えることと関連していた。これらの結果は、事前学習済みモデルの周辺を勾配なしで効率よく探索する際に、摂動の幾何学的な設計が重要であることを示す。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

RandOpt samples weight-perturbed language models and ensembles top-ranked candidates through plurality voting, but its global perturbation scale ignores heterogeneous module geometry. We propose Modular Norm RandOpt, an architecture-aware sampling method using module-wise natural norms and calibrated scales while preserving selection and voting. It outperforms RandOpt using $3\times$ fewer candidates on Countdown and at least $12\times$ fewer on GSM8K, with corresponding wall-clock savings. Evaluations across seven tasks and three Qwen scales ($0.5$B--$3$B) show higher mean accuracy than RandOpt on Countdown, GSM8K, and MATH-500 at every scale. The gains extend to Llama 3.2 $3$B and Gemma 3 $4$B on Countdown and GSM8K. On Qwen2.5-1.5B, our ensembles also achieve higher mean accuracy than iterative baselines on both tasks at comparable main-run evaluation budgets. On GSM8K, a tail-density diagnostic implies only a $1.2$--$1.8\times$ candidate reduction, while most ensemble improvement is associated with more favorable correct-expert support. These results highlight perturbation geometry as a key design choice for population-efficient, gradient-free search around pretrained models.

著者のコメント

Preprint. Project page: https://kiratoyoshihara.github.io/Modular-Norm-RandOpt-page/

arXiv ID: 2609.25745 / 要約の誤りについて