反発する自己注意に現れるカオスと注意の集中
Nonequilibrium Phases of Repulsive Self-Attention: Chaos, Attention Condensation, and Emergent Locality
この論文をやさしく読む
ひとことで言うと
反発する単純な自己注意モデルで、周期運動、カオス、注意の集中がどう生じるかを調べた。
何に役立つ?
考えられる用途は、再帰型注意機構の力学的な理解である。要旨は最小モデルの理論解析とシミュレーションを述べる。
この研究の面白いところ
二次元では注意の集中にβ∼N²が必要だが、高次元ではβが定数でも集中の証拠が得られた。
どこまで分かった?
高次元の転移はシミュレーションによる証拠として述べられる。実際の大規模言語モデルで同じ相が現れるとは示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本研究は、正規化されたN個のトークンを持ち、Q=K=I、値の写像V=−Iとした最小限の再帰型トランスフォーマーの非平衡ダイナミクスを調べる。類似度に基づく注意は近い表現を選ぶ一方、負の値写像がトークンを選ばれた場から遠ざける。このフィードバックにより、表現の幾何構造と注意ネットワークの双方が絶えず組み替わり得る。 次元d=2ではトークンは円上にあり、正多角形が厳密な固定点となる。注意のフィードバック強度γを上げると、この多角形は反転分岐によって安定性を失い、周期2の運動、カオス、クラスタの交換や反転状態が現れる。この時間的な複雑さにもかかわらず、ソフトマックスの鋭さβを有限の固定値に保ちNを無限大にすると、注意は広く分散したままである。注意の集中は、βがN²に比例するスケーリング領域で初めて現れる。経路選択が硬い極限では、反発する更新が局所的な摂動を増幅し、経路の相手の切り替わりがそれを弾道的に伝え、表現空間に「バタフライ・コーン」が生じる。 高次元の幾何は、注意が局所化する別の道を与える。d=Nとして無限大に近づける場合、ガウス分布に従う初期条件からのシミュレーションは、動的に生じる有限の重なりの隔たりによって、β=O(1)で集中への転移が起こる証拠を示す。γによって、分散した単体状の状態、全体でそろった反転、カオスの兆候を伴う集中した能動的な経路選択、分裂したクラスタの反転などの相が現れる。これらは、時間的な活動、注意の集中、幾何学的なクラスタ化が別々の集団現象であり、疎な注意でも動きが凍結せず持続し得ることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We study the nonequilibrium dynamics of a minimal recurrent transformer with $N$ normalized tokens, $Q=K=I$, and a negative value map $V=-I$. Similarity-based attention selects nearby representations, while the negative value map drives tokens away from the selected field. This feedback can continually reorganize both the representation geometry and the attention network. For $d=2$, the tokens lie on a circle, where the regular polygon is an exact fixed point. As the attention feedback strength $\gamma$ is increased, the polygon loses stability through a flip bifurcation, giving rise to period-two motion, chaos, and cluster-exchange or cluster-flip states. Despite this temporal complexity, attention remains diffuse as $N\to\infty$ at finite fixed softmax sharpness $\beta$. Attention condensation instead emerges in the scaling regime $\beta\sim N^2$. In the hard-routing limit, repulsive updates amplify local perturbations and routing-partner switches transmit them ballistically, producing an emergent butterfly cone in representation space. High-dimensional geometry provides a distinct route to localization. For $d=N\to\infty$, simulations from Gaussian initial conditions provide evidence for a condensation transition at $\beta=O(1)$, driven by dynamically generated finite overlap gaps. Depending on $\gamma$, the resulting phases include diffuse simplex-like states, consensus flips, condensed active routing with signatures of chaos, and fragmented cluster flips. These results establish temporal activity, attention condensation, and geometric clustering as distinct collective phenomena, and show that sparse attention can sustain persistent dynamics rather than freeze it.
著者のコメント
54 pages, 21 figures, including appendices
arXiv ID: 2609.28448 / 要約の誤りについて