ニューラルネットは状態そのものか剰余類かを追跡する
Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking
この論文をやさしく読む
ひとことで言うと
群の要素を順番に掛け合わせる課題で、モデルが正確な状態の代わりに剰余類だけを覚えている場合を調べた研究です。
何に役立つ?
系列を処理するモデルの正解率から内部で保持する情報を読み解くための理論と実験結果です。要旨に一般的な実用課題での性能改善は示されていません。
この研究の面白いところ
Transformerと再帰型ネットワークでは、訓練中に追跡できる剰余類の種類が異なりました。再帰状態の一部を入れ替える実験でも、剰余類の情報が移ることを示しています。
どこまで分かった?
理論証明は有限群と一様な独立同分布入力などの条件の下で述べられています。ほかの入力分布や課題への一般化は要旨からは分かりません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
状態追跡には一連の更新の合成が必要だが、予測の正解率だけではモデルが何を学んだかは分からない。本研究は、群の要素の累積積を予測するよう訓練したニューラルネットワークを調べる。Transformerでは、商類を復元しながら、その類に属する要素の間ではほぼ一様に予測する商に基づく解を見いだした。類の大きさの逆数は、当てはめたパラメータなしで部分的な正解率を予測し、パリティに基づく説明をそれ以外の商へ広げる。基準となるTransformerの予測は、正確な追跡ができる範囲を超えると、入力接頭辞の順序を入れ替えてもほとんど変わらない。 有限群で、群全体からの一様な独立同分布入力を仮定すると、順序を考慮しない場合の最適な厳密正解率は、接頭辞が長くなるにつれてアーベル化の類の大きさの逆数へ収束することを証明し、観測されたアーベル化の水準と整合する。一方、逐次更新ならより多くの情報を保てる。正規部分群かどうかにかかわらず、部分群の右剰余類による任意の分割は逐次更新を経ても維持できる。標準的なTransformerの調査では、復元された剰余類の分割はすべて正規部分群に由来したが、パラメータ数をそろえた再帰型ネットワークは訓練中に正規・非正規の右剰余類を追跡する段階を経た。A₅では、非正規剰余類を符号化する再帰状態の低次元部分空間を特定した。三次元の場合、剰余類の平均ベクトルは近似的な十二面体を形作り、その部分空間の状態成分を入れ替えると、共通の後続入力を通して提供元の剰余類状態が移る。結果は、モデルが追跡する部分群の剰余類を通じて、部分的な正解率、学習の段階、内部計算を結び付ける。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
State tracking requires composing a sequence of updates, but accuracy alone does not reveal what a model has learned. We study neural networks trained to predict the running product of group elements. We identify quotient solutions in Transformers, where models recover the quotient class while predicting nearly uniformly among its members. The reciprocal of class size predicts partial accuracy without a fitted parameter, extending parity-based accounts to non-parity quotients. Our baseline Transformers' predictions change little under prefix reordering beyond the exact-tracking frontier. We prove that, for finite groups under uniform i.i.d. full-group inputs, optimal order-blind exact accuracy converges to the reciprocal of abelianization class size as prefix length grows, consistent with the observed abelianization plateaus. Sequential updates permit more: any partition into right cosets of a subgroup, normal or not, survives sequential updates. In our census of standard Transformers, every recovered coset partition comes from a normal subgroup, whereas parameter-matched recurrent networks pass through both normal and non-normal right-coset stages during training. On $A_5$, we identify low-dimensional subspaces of the recurrent state that encode non-normal cosets. In the three-dimensional cases, coset mean vectors form approximate dodecahedra, and swapping the state components in these subspaces transfers the donor's coset state through a shared input suffix. Our results connect partial accuracy, learning stages, and internal computation through the subgroup cosets that models learn to track.
著者のコメント
69 pages including appendices; 9 pages of main text
arXiv ID: 2609.29951 / 要約の誤りについて