arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

残したい概念を守りながら画像生成モデルから対象を消す

RASteer: Retain-Aware Activation Steering for Concept Erasure in Diffusion Models

Yongliang Wu, Haori Lu, Yulun Wu, Jinqi Luo, Xingyu Zhu, Yaoyao Liu

この論文をやさしく読む

ひとことで言うと

画像生成AIから特定の概念を消すとき、残したい概念と共通する内部表現を見分けて、消し過ぎを抑える方法です。

何に役立つ?

特定の作風や対象を生成しにくくしつつ、無関係な画像の生成能力を保ちたい場合に使うことが考えられます。追加学習なしで推論時の活性を調整します。

この研究の面白いところ

共有部分を一律に保護すると消去が弱くなるため、層とノイズ除去ステップごとに重なりを測って保護量を調整します。

どこまで分かった?

優位性は評価対象のモデルとベンチマークでの比較に限られ、要旨には具体的な数値や全モデルの名称はありません。対象の完全な消去やあらゆる保持概念の保全を保証したとは述べていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

概念消去は、著作権のある作風、識別可能なキャラクター、安全上問題のあるコンテンツなどの対象概念を、学習済みのテキストから画像への拡散モデルから取り除きつつ、他のコンテンツを生成する能力を保つことを目指す。既存の活性操作手法は、主に対象概念から消去方向を作り、推論時にその方向に沿ってモデルの活性を調整する。しかし、対象概念と保持したい概念は、モデルの表現空間で重なり合うことが多い。そのため、この方向には保持対象の概念が依存する共有成分も含まれる。この方向に沿って直接操作すると、保持したい概念も抑制し、対象外のコンテンツの生成を損なう可能性がある。 この問題に対し、本研究では学習不要の手法Retain-aware Activation Steering(RASteer)を提案する。RASteerはまず、保つべき概念から保持部分空間を作る。次にRetain-Orthogonal Steering(ROS)が、この部分空間に沿った成分を消去方向から取り除き、対象に特化した操作にする。ただし、共有成分を完全に除くと消去が弱まる可能性があるため、さらにOverlap-Adaptive Calibration(OAC)を導入する。OACは各層と各ノイズ除去ステップで、消去方向と保持部分空間の重なりを使って各共有成分をどれだけ取り除くかを制御し、対象の消去と概念の保持を両立させる。 複数の基盤モデルとベンチマークにおける、安全上問題のあるコンテンツ、個別対象、芸術的な作風の消去実験では、RASteerは評価した活性操作および重み編集の基準手法と同等以上の性能を示し、消去と保持のよりよい均衡を達成する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-image diffusion model while preserving its ability to generate other content. Existing activation steering methods build an erasure direction mainly from the target concept and adjust model activations along it at inference time. However, target and retained concepts often overlap in the model's representation space, so this direction also contains shared components that retained concepts rely on. Steering directly along this direction can therefore suppress retained concepts and harm the generation of non-target content. To address this issue, we propose Retain-aware Activation Steering (RASteer), a training-free method. RASteer first builds a retain subspace from the concepts to preserve. Retain-Orthogonal Steering (ROS) then removes components aligned with this subspace from the erasure direction, making steering more specific to the target. Since fully removing the shared components can weaken erasure, we further introduce Overlap-Adaptive Calibration (OAC). At each layer and denoising step, OAC uses the overlap between the erasure direction and the retain subspace to control how much of each shared component is removed, balancing target erasure and concept preservation. Experiments on unsafe-content, instance, and artistic-style erasure across multiple backbones and benchmarks show that RASteer matches or outperforms the activation steering and weight editing baselines we evaluate, achieving a better balance between erasure and preservation.

著者のコメント

20 pages. Project page: https://rasteer.cvmlgroup.web.illinois.edu/

arXiv ID: 2610.01969 / 要約の誤りについて