言語モデル内部の非線形な概念構造を層・モデル間で比べる
Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models
この論文をやさしく読む
ひとことで言うと
言語モデルの内部概念を直線的な方向だけで表さず、曲がった低次元の構造として抽出し、層やモデル、学習段階ごとに比べています。
何に役立つ?
モデル内部の表現が学習や微調整でどの程度変わるかを調べる解析手段になります。多言語間や同系列モデル間で、概念がどれほど共有されるかを評価する用途もあります。
この研究の面白いところ
線形の比較指標では見えにくかった層のブロック構造を検出しています。また、調べたTulu-3ではSFTで大きな変化が起こり、その後の選好学習では初期層がほぼ維持されるという差を示します。
どこまで分かった?
知見は要旨に挙げられたモデルと学習段階での解析です。GPT-2での概念共有の不検出を、あらゆる多言語知識の不存在と読み替えることはできません。幾何学的な整合度は、それ自体で振る舞いの因果説明を確定するものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)の情報処理を理解するには、内部のトークン表現が幾何学的にどう組織されているかを分析する必要がある。既存の機構的解釈可能性(MI)の手法は概念の抽出を目指すが、強い線形性の仮定に制約されており、非線形な特徴多様体の証拠がこの仮定に疑問を投げかけている。本研究では、コンピュータービジョンの非線形多次元概念発見(NLMCD)をLLMのトークン単位の活性へ適応し、概念を低次元多様体としてモデル化することで、線形概念を超える。層やモデルをまたいで概念多様体を比較するため、明示的に特徴を対応付けず幾何学的な近さを測る、一般化Rand指数である概念ベース整合度(CBA)スコアを導入する。 解析から六つの主要な知見を得た。(i)隣接層を使った妥当性確認では、CBAはPCAやCKAに基づく線形の比較指標より感度が高い。(ii)層ごとの整合度行列は中間層と後段層に二つのブロック構造を示し、モデル間で一貫しているが、線形指標では見えにくい。(iii)ネットワークの大部分では概念の構成は構文が支配的で、後段層になると構文と意味が混ざった概念が増え、最終層へ向かって出力志向が強まる。(iv)英語と中国語の概念共有は普遍的ではなく学習に依存し、Qwenで最も強く、Llamaでは弱く、GPT-2には見られない。(v)モデル間の整合もこの構造を反映し、規模の異なる同系列のQwenモデル間では強い対応がある一方、モデル系列をまたぐ整合は弱い。(vi)Tulu-3の学習段階間では隣り合う段階で整合度が最も高く、最大の変化はベースモデルから教師あり微調整(SFT)への移行で起こる。その後の選好アラインメント段階(DPO、RLVR)では初期層はほぼ変わらず、RLVRは後段層でもDPOの概念をおおむね保持する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Understanding information processing in large language models (LLMs) requires dissecting the geometric organization of their internal token representations. While existing mechanistic interpretability (MI) methods seek to extract concepts, they are constrained by a strong linearity assumption challenged by evidence of non-linear feature manifolds. We move beyond linear concepts by adapting Non-Linear Multi-Dimensional Concept Discovery (NLMCD) from computer vision to token-level LLM activations, modeling concepts as low-dimensional manifolds. To compare concept manifolds across layers and models, we introduce a concept-based alignment (CBA) score, a generalized Rand index that measures geometric proximity without explicit feature matching. Our analysis yields six key findings: (i) a neighboring-layer sanity check shows CBA is more sensitive than PCA- or CKA-based linear baselines; (ii) layer-by-layer alignment matrices reveal two block structures in intermediate and late layers, consistent across models and obscured by linear metrics; (iii) concept composition remains syntax-dominated through most of the network before giving way to increasingly mixed syntactic-semantic concepts in later layers, with increasing output-orientation toward the final layers; (iv) multilingual concept sharing between English and Mandarin is training-dependent rather than universal, strongest in Qwen, weaker in Llama, and absent in GPT-2; (v) inter-model alignment mirrors this structure, with strong correspondence between same-family Qwen models of different scale but weak alignment across model families; and (vi) across Tulu-3 training stages, alignment is highest between adjacent stages, with the largest shift between the base model and SFT, while subsequent preference-alignment stages (DPO, RLVR) leave early layers largely unchanged and RLVR mostly preserves DPO's concepts in late layers.
著者のコメント
24 pages, 13 figures. Code: https://anonymous.4open.science/r/NLMCD-NLP-C5E7
arXiv ID: 2610.01821 / 要約の誤りについて