共有埋め込みを持つモデルをAdamWなしで学習する
AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models
この論文をやさしく読む
ひとことで言うと
入出力で同じ埋め込み表を共有するモデルに対し、MuonとAdamWを併用せず、二次モーメントを保存しない方式で全パラメータを学習する方法です。
何に役立つ?
モデル学習の最適化器が使うメモリを減らすために役立ちます。9設定のベンチマークで、Hybrid Muon比で状態メモリを約20%削減し、平均検証損失とパープレキシティも改善したと報告しています。
この研究の面白いところ
共有語彙テーブルには入力側の疎な勾配と出力側の密な勾配が届くため、普通の行列と同じ更新では扱いにくい点に着目しています。有限の上限を設けて、大きさの情報保持と更新の集中の抑制を両立させています。
どこまで分かった?
約20%は最適化器の状態メモリの削減で、学習全体のメモリ削減率ではありません。時間増加約1%は条件をそろえた学習での値です。要旨の言語モデル規模は1億2400万〜10億パラメータで、それを超える規模の結果は記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Muonは、行列パラメータにスペクトルノルムに関する最急降下更新を適用することで、大規模学習を改善する。しかし、実際のモデルには、密行列の幾何に適合しないパラメータブロックも含まれる。重要な例が、言語モデルなどのトークン生成器に現れる、入出力で共有された語彙テーブルである。このテーブルには、疎な入力参照から密な出力分類器の更新まで、構造が異なる複数の源から勾配が届きうる。標準的な構成では、こうしたブロックは補助的なAdamW最適化器に任される。そのため二次モーメントの状態が再び必要となり、別名で参照される同じテーブルが一般的なテンソルとして更新される。 本研究では、MuonをAdamWなしで拡張するAF-Muonを提案する。隠れ層の重み行列にはMuonの行列更新を保持しつつ、共有語彙テーブルには非ゼロ要素の支持を考慮した有限上限付き線形最小化オラクルを用い、1次元の補助パラメータにはRMSで正規化した更新を用いる。したがってAF-Muonは、すべてのパラメータ種別を、一次モーメントのバッファ1つだけで、二次モーメントの状態なしに学習する。本研究のベンチマークでは、Hybrid Muonに比べて最適化器の状態メモリを約20%削減する。 テキスト、画像、タンパク質配列データにまたがる9種類のトークン共有設定で、AF-MuonはHybrid MuonとSCION型Sign方式の両方より、平均検証損失とパープレキシティを改善する。対象には、1億2400万〜10億パラメータのデコーダ専用言語モデル、完全共有のT5型エンコーダ・デコーダ、ImageGPT型の画像トークンモデル、タンパク質モデル、疎なMoEの派生モデルが含まれる。 長期間の学習とハイパーパラメータ感度の調査により、改善が頑健であることを確認する。また、同じモメンタムを使った診断から、改善は有限上限に由来することを示す。この上限は、Signより行内の大きさの情報を多く保つ一方で、行RMSの座標への集中を制限する。これらの結果は、共有語彙テーブルを独自の最適化幾何として位置付け、モデル、モダリティ、構造をまたいで頑健な、AdamWを使わないMuonの方式を与える。条件をそろえた学習での1ステップ当たりの時間増加は約1%である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Muon improves large-scale training by applying a spectral-norm steepest-descent update to matrix parameters, but practical models also contain parameter blocks that do not fit dense-matrix geometry. One important case is the tied vocabulary table, which appears in language models and other token generators and can receive multiple structurally different gradient sources, from sparse input lookups to dense output-classifier updates. In the reference recipe these blocks are handed to an auxiliary AdamW optimizer, which restores second-moment state and updates the aliased table as a generic tensor. We propose AF-Muon, an AdamW-free extension of Muon that keeps the Muon matrix update for hidden weight matrices while using a support-aware finite-cap linear minimization oracle for tied vocabulary tables and an RMS-normalized update for one-dimensional auxiliary parameters. AF-Muon therefore trains every parameter class with a single first-moment buffer and no second-moment state, saving around 20% optimizer-state memory relative to Hybrid Muon in our benchmark. Across nine tied-token settings - decoder-only language models from 124M to 1B parameters, a fully shared T5-style encoder-decoder, and ImageGPT-style image-token, protein, and sparse-MoE variants, spanning text, image, and protein-sequence data - AF-Muon improves mean validation loss and perplexity over both Hybrid Muon and a SCION-style Sign endpoint. Long-horizon runs and hyperparameter sensitivity studies confirm the gain is robust, and identical-momentum diagnostics attribute it to the finite cap, which preserves more within-row magnitude than Sign while bounding the coordinate concentration of row-RMS. These results identify tied vocabulary tables as a distinct optimizer geometry and yield a robust AdamW-free Muon variant across models, modalities, and architectures, with about 1% step-time overhead in matched training.
著者のコメント
63 pages, 19 figures. An earlier, shorter version of this work was accepted as a poster at the OPT 2026 workshop (Optimization for Machine Learning) at NeurIPS 2026; this is the complete version
arXiv ID: 2610.01395 / 要約の誤りについて