勾配で方向を決め、関数評価で更新幅を選ぶ追加学習
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
この論文をやさしく読む
ひとことで言うと
モデルの重みをどちらへ動かすかは通常の勾配法に任せ、どれだけ動かすかを追加の関数評価で決めます。方向と更新幅を別々に扱う方法です。
何に役立つ?
言語モデルの追加学習で更新幅を適応させる設計に役立ちます。固定幅のベースラインに対する最適化や最終性能の改善が複数の設定で報告されています。
この研究の面白いところ
現在の勾配と追加2回の評価から、選んだ方向の曲率を推定します。大掛かりな直線探索を行わずに、局所モデルで更新幅を選ぶ点が特徴です。
どこまで分かった?
保証される収束先は停留点の近傍です。実験の改善はすべての設定で一律ではなく、目的関数により幅や適切な局所モデルが変わります。理論の詳細な仮定と各実験の数値は要旨にはありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模ニューラルネットワークの最適化では、ステップ幅の選択が依然として中心的な課題である。慎重すぎる更新は収束を遅らせ、積極的すぎる更新は不安定化させる。私たちは零次・一次の最適化を組み合わせるZFO(Zero-and-First-Order optimization)により、方向の選択とステップ幅を分離する軽量な枠組みを提案する。ZFOは信頼できる一次の最適化器で方向を決め、その1次元部分空間に沿ってだけ零次の評価を行い、どこまで進むかを選ぶ。現在の勾配情報と追加2回の目的関数評価を使い、提案された方向に沿った目的関数の局所モデルを構築して、有界な探索区間内で曲率を考慮したステップを選択する。これにより、完全な直線探索より低コストな適応的ステップ選択が得られる。 理論的な保証として、サンプルを共有する評価から信頼できる有限差分の曲率推定が得られること、得られた局所モデルが探索区間でほぼ最適なステップを選ぶこと、ZFOが停留点の近傍に収束することを示す。評価した設定、言語モデル、データセットでは、ZFOは固定ステップの一次手法のベースラインに対し、最適化と最終性能を改善することが多かった。ただし改善の大きさと適した局所モデルは目的関数に依存する。コードはhttps://github.com/nizswan/Zeroth-First-Order-Framework で公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 掲載先の記載あり
著者による掲載先の記載:40th Conference on Neural Information Processing Systems (NeurIPS 2026)。出版社での独立確認は未実施です。
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine \textbf{Z}ero-and-\textbf{F}irst-\textbf{O}rder optimization~(ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current {gradient information} and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less than a full line search. We provide theoretical guarantees to show that shared-sample evaluations produce reliable finite-difference curvature estimates, that the induced local model selects a near-optimal step along the search interval, and that ZFO converges to a neighborhood of a stationary point. Across the evaluated settings, language models and datasets, ZFO frequently improves optimization and final performance relative to fixed-step first-order baselines, with the magnitude and preferred local model depending on the objective. Our code is publicly available at: https://github.com/nizswan/Zeroth-First-Order-Framework.
著者のコメント
Accepted to 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Code: https://github.com/nizswan/Zeroth-First-Order-Framework
arXiv ID: 2610.02190 / 要約の誤りについて