arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

ニューラルネットが特徴を削除せず隠す仕組み

Hidden not Deleted: How Networks Suppress Entangled Features

Akash Samanta, Manish Pratap Singh, Debasis Chaudhuri

この論文をやさしく読む

ひとことで言うと

二つの特徴が重なって表現されると、ネットワークは標的を消す代わりに隠してしまう場合がある。

何に役立つ?

モデルから概念を除く手法やアンラーニングの評価に役立つ。

この研究の面白いところ

初期値に応じた二つの回路解を見つけ、隠れた特徴が一つの値の差し替えで戻ることを示した。

どこまで分かった?

構成した特徴の絡み合いとネットワークでの機構的な結果である。全ての大規模言語モデルで同じ仕組みが働くとは示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

線形射影による概念消去法は、特徴が分離可能な部分空間にあると仮定する。本研究は、特徴が密に重なって表現される場合、この仮定が崩れることを示す。二つの特徴を一つの部分空間を共有する正反対の組として配置すると、最先端の線形消去法は標的だけでなく両方を壊してしまう。勾配降下で学習したネットワークは、この問題を非線形に解くが、その方法は一様ではない。初期値に応じて、鏡像解と影解と名付けた、回路レベルで異なる二つの解のいずれかに収束する。 特徴の絡み合いの度合いに対してこの分岐を描き、実験設定の偶然ではなく安定したアトラクター構造を反映することを示す。さらに、狙いを定めた因果的介入によって、両方の解で消去されたはずの特徴表現が、測定可能な大きさで残ることを示した。その表現は追加学習なしに、単一のスカラー値の差し替えで回復できた。これは、忘却させた知識が削除されず抑制されるために再び現れるという、大規模言語モデルのアンラーニングで最近実験的に観測された失敗に似ている。本研究は、その失敗が起こる理由について、仕組みに基づき因果的に検証した説明を与える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Concept erasure methods that operate via linear projection assume that features occupy separable subspaces. We show this assumption fails under dense superposition: when two features are forced into an antipodal pair sharing a single subspace, state-of-the-art linear erasure destroys both, not just the target. Networks trained with gradient descent instead solve this problem non-linearly, but not uniformly: they converge to one of two distinct circuit-level solutions depending on initialization, which we call mirror and shadow solutions. We map this bifurcation as a function of feature entanglement, show it reflects a stable attractor structure rather than an artifact of our setup, and use targeted causal interventions to demonstrate that both solutions leave a substantial, measurable trace of the erased feature's representation intact, recoverable through a single scalar patch rather than requiring any further training. This mirrors a failure mode recently observed empirically in LLM unlearning, where suppression rather than deletion allows forgotten knowledge to resurface; our results offer a mechanistic, causally-validated account of why that failure mode occurs.

arXiv ID: 2609.27593 / 要約の誤りについて