線形表現仮説には表現の同値性の指定が必要
The Linear Representation Hypothesis Needs a Group Action
この論文をやさしく読む
ひとことで言うと
AIモデルの内部表現が線形だという主張は、どの変換後の表現を同じと見なすかで意味が変わると整理した。
何に役立つ?
モデルの解釈可能性研究で、指標や介入が実際にどの仮説を検証しているか確認する枠組みになる。
この研究の面白いところ
表現、生成手順、主張する性質を群作用で形式化し、見かけ上同じ分析の前提の違いを明らかにする。
どこまで分かった?
要旨が示すのは概念の形式化と既存分析の点検であり、特定モデルで線形表現仮説が正しいと実証する結果ではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
特定の学習済みモデルを超えて表現について主張するには、二つの表現をいつ同等と見なすかを定める必要がある。線形表現仮説は、この同値性を明示せずに議論されることが多い。同値性の定義が違えば保たれる構造も違うため、同じ表現を調べているように見える指標、プローブ、介入が、実際には別々の仮説に対応し得る。本研究は、線形表現仮説は単一の仮説ではなく、表現の同値性によって区別される主張の一群だと論じる。この考えを群作用を用いて形式化し、表現という対象、それを作り出す手順、最終的に主張する性質を指定するとともに、モデルの構造が課す同値性も考慮する。この枠組みにより、指標、表現を読み出す位置、分析の段階によって前提がどう変わるかを明確にし、一般的な表現の量と最近の解釈可能性分析を点検する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
To make claims about representations that generalize beyond a particular trained model, we need to specify when two representations should count as equivalent. The Linear Representation Hypothesis is often discussed without making this equivalence explicit. Different notions of equivalence preserve different structures, so metrics, probes, and interventions that appear to study the same representation may in fact correspond to different hypotheses. We therefore argue that the Linear Representation Hypothesis is not one hypothesis but a family of claims distinguished by representation equivalence. We formalize this idea using group actions, specifying the representation object, the procedure that produces it, and the property ultimately asserted, while accounting for equivalences imposed by the model architecture. This framework clarifies how assumptions can change across metrics, reading points, and analysis stages, and we use it to audit common representation quantities and recent interpretability analyses.
著者のコメント
16 pages, 1 table
arXiv ID: 2609.27158 / 要約の誤りについて