arXiv論文メモ
新着一覧
cs.CL / cs.AI / cs.LG · 査読状況未確認

大規模言語モデルRufus-Airの公開追加学習手順

Rufus-Air: An Open LLM Post-Training Recipe

Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He, Hongye Jin, Linwei Li, Shiyang Li, Fenglin Liu, Xin Liu, Priyanka Nigam, Haoyang Wen, Zhenghao Xu, Zhuocheng Xu, Bing Yin, Qingyu Yin, Chao Zhang, Rongzhi Zhang, Zhihan Zhang, Zixuan Zhang, Zixuan Zhang, and Tuo Zhao

この論文をやさしく読む

ひとことで言うと

大規模言語モデルの追加学習を8段階で行う公開手順と、その段階ごとの知見。

何に役立つ?

追加学習の再現や、データ・報酬・学習段階の設計を比較する材料になる。

この研究の面白いところ

報酬の信頼性を段階の順序に結び付け、設備や実装上の選択も再現可能な手順に含めた。

どこまで分かった?

要旨は公式版との改善と同規模の公開モデルとの競争力を述べるが、個別の評価値は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Rufus-AirはGLM-4.5-Air-Base(106B-A12B)を対象とした、公開され再現可能な追加学習手順である。教師あり微調整、推論の強化学習、コードの強化学習、指示追従の強化学習、汎用エージェント、コーディングエージェント、検索エージェント、RLHFという8段階の直列パイプラインからなる。再現に必要なデータ、報酬設計、基盤設備、段階の順序、各段階の結果を記録する。学習は基本能力から高度な能力へ、検証しやすい明確な報酬から判定者に基づく柔らかい信号へ進む。公開ソフトウェアと公開データを基盤とし、多くは公開時の形のまま使用し、新たな人手による注釈や内部の蒸留教師は使わない。主な知見は、多様で質の高い教師あり微調整が能力の強い土台となること、難易度による選別が強化学習の問題を有効な学習範囲に保つこと、報酬の信頼性が段階を並べる実用的な原則になること、設備やエンジニアリング上の選択も手順の一部であることの四つである。Rufus-Airは公式のGLM-4.5-Air追加学習版を上回り、同程度の規模の公開モデルと競争力がある。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Rufus-Air is an open and reproducible post-training recipe on GLM-4.5-Air-Base (106B-A12B), organized as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction-Following RL, General Agent, Coding Agent, Search Agent, and RLHF. We document the data, reward design, infrastructure, stage order, and stagewise results needed to reproduce the recipe. Stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge-based signals. Training builds on open-source components and public data, much of it used as released, without new human annotation or an in-house distillation teacher. Our main findings are that (i) diverse, high-quality SFT establishes a strong capability floor; (ii) difficulty filtering keeps RL prompts within a productive learning range; (iii) reward reliability provides a practical principle for ordering stages; and (iv) infrastructure and engineering choices are part of the recipe, not just an implementation detail. Rufus-Air improves over the official GLM-4.5-Air post-trained release and is competitive with similarly sized open models.

著者のコメント

47 pages, 9 figures, 20 tables. Authors are listed alphabetically by surname; all contributed while at Amazon. The two authors named Zixuan Zhang are different people

arXiv ID: 2609.29421 / 要約の誤りについて