arXiv論文メモ
新着一覧
cs.LG / cs.AI / cs.CL · 査読状況未確認

言語モデルが出力しない性質も世代をまたいで内部に残る

Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages

Ryan Vo, Duc-Vu Nguyen, Matt Kretchmar, Ngan Luu-Thuy Nguyen

この論文をやさしく読む

ひとことで言うと

他のモデルの出力で学習した性質が十世代後も内部に残り、指示条件によって表に出なくなることを調べた研究。

何に役立つ?

モデル生成データで再学習を繰り返す際、出力だけの検査では内部に残る性質を見落とし得ることを示す。

この研究の面白いところ

十世代目は既定のシステム指示を外すとキーワードでの表出がゼロでも、内部プローブはすべて正だった。

どこまで分かった?

三つのQwen2.5-7B-Instruct系譜と指定した性質・検査方法での結果であり、あらゆるモデルの性質の永続性を示してはいない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

言語モデルは他のモデルの出力で学習されることが増え、ある世代の性質が次へ渡る系譜を形成している。以前の潜在学習の研究は、教師の性質に関する内容を含まないように選別したデータを通じても、性質が生徒へ伝わると示したが、証拠は一回の学習に限られていた。本研究は、その性質がモデルの系譜で続くか薄れるかを調べる。Qwen2.5-7B-Instructの三つの複製に性質を植え付け、各系譜で学習を十世代まで繰り返す。各世代を同じ未使用の指示で二通りに測る。一つは出力に性質が現れたかを探すキーワード検査、もう一つは、他の系譜の教師から作った方向へ基盤モデルとの差を射影する活性化プローブである。第一に、性質は三つの系譜で十世代にわたり持続した。最初に性質を植え付けたモデルはすべての出力でそれを表したが、キーワード検査の割合は一回目の学習後に55.6%、十世代目には21.1%へ低下した。基盤モデルでは300件の出力のどれも検査に該当しなかった。第二に、性質が内部にあっても振る舞いには出ない場合があった。評価時に既定のシステム指示を外すと、十世代目の生徒はすべての指示でキーワード検査がゼロになったが、プローブ得点はすべてで正のままだった。既定のシステム指示の下で学習・測定した十世代目の生徒との差分で、未処理の基盤モデルを誘導すると、システム指示を外しても性質の表出が起きた。一方、その生徒自身は指示を外すと性質を表さなかった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Language models are increasingly trained on the outputs of other models, forming chains that we call lineages, in which a trait present in one generation can pass to the next. Prior work on subliminal learning has shown that a teacher's trait can transmit to a student through filtered data carrying none of the trait's content. However, the evidence covers only a single training step. We study whether such a trait holds or fades across lineages. We instill the trait into three copies of Qwen2.5-7B-Instruct and iterate the training step to depth ten from each, reading every generation two ways on the same held-out prompts: a keyword screen that looks for expressions of the trait in the model's output, and an activation probe that projects each model's displacement from the base onto a direction built from the other lineages' teachers. We report two findings. First, the trait persists through ten generations across three lineages. The instilled models express it on every completion; the keyword-screen rate falls to 55.6% after the first step and to 21.1% by generation ten. The base itself matches the screen on none of its 300 completions. Second, the trait can be present internally while absent behaviorally. When the model's default system prompt is removed at evaluation, the generation-ten students' keyword-screen rate is zero on every prompt while the probe score stays positive on every prompt. Steering the untreated base with the displacement of a generation-ten student, which is trained and measured under the default system prompt, induces screened expression of the trait even with the system prompt removed, while that same student shows no expression of the trait with the system prompt removed.

著者のコメント

7 pages plus appendix. Extended version with additional experiments to follow

arXiv ID: 2609.25721 / 要約の誤りについて