arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

アラインメント中間学習の有効性を厳しい条件で検証する

Stress-testing Alignment Midtraining

Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan

この論文をやさしく読む

ひとことで言うと

AIに安全性関連の文書を中間学習させると、後の学習で望ましい振る舞いが広く身に付くかを検証します。

何に役立つ?

学習時に見せていない状況への振る舞いの一般化を、安全性手法の効果としてどこまで期待できるか判断する材料になります。

この研究の面白いところ

単純な設定で誘導できた動機が、ごく少量の競合する微調整データで消えることを報告します。規則の説明だけでなく、どこかの学習段階で実例が必要だった点も特徴です。

どこまで分かった?

最大1100億パラメータ、10億中間学習トークンまで評価しています。著者の結論は公開証拠が十分でないというもので、中間学習が全設定で無効だと証明したものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

最先端モデルを事後学習の手法でアラインメントする際、考えられるすべての運用環境でモデルに示してほしいすべての振る舞いを、直接例示することはできない。モデルは事後学習の分布の外へ汎化する必要がある。提案されている解決策の一つが、アラインメント中間学習(AMT)である。これは、後の学習段階での汎化を促すため、アラインメントに関連する大量の文書を使って事前学習を継続する方法である。AMTはアラインメントの方法として注目されているが、その有効性を示す公開された証拠は限られている。 この問題に対処するため、中間学習に関するいくつかの仮定を特定し、最大1,100億パラメータのモデルと10億中間学習トークンまで、規模を変えて評価する。例えば、事後学習データが2つの考えられる動機のどちらとも解釈できる場面を研究する。この設定の単純な場合には、中間学習でモデルの動機を方向付けられる。しかし、競合する動機を示唆する微調整データがごくわずかな割合でも含まれると、AMTの効果は消失する。 また、AIに複数の規則を守らせたいが、その一部だけを例示する場面も研究する。これらの規則を頑健に学習させるには、中間学習または事後学習のデータセットのどちらかに、例示が存在しなければならないことが分かった。これらを含む結果に基づき、強力なAIシステムのアラインメントに本質的に伴う中心的な困難に中間学習で対処できると確信を持って述べるには、公開された証拠は不十分であると著者らは考える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribution. One proposed solution is alignment midtraining (AMT), which continues pretraining on large volumes of alignment-relevant documents to encourage generalisation in later stages of training. Despite the prominence of AMT as an alignment approach, there is limited public evidence for its effectiveness. To resolve this, we identify several assumptions around midtraining and evaluate them across scale: up to 110 billion-parameter models and 1 billion midtraining tokens. For instance, we study a scenario where post-training data is ambiguous between two possible motivations. We find that midtraining can steer the model's motivation in simple versions of this setting. However, the presence of a tiny fraction of finetuning data which suggests a competing motivation erases the effects of AMT. We also study scenarios in which we want an AI to follow a number of rules, but only demonstrate a subset of them. We find that demonstrations must be present either in midtraining or post-training datasets for these rules to be robustly learned. Based on these and other findings, we do not believe that there is sufficient public evidence for us to confidently state that midtraining can address the core difficulties inherent in aligning powerful AI systems.

arXiv ID: 2609.20412 / 要約の誤りについて