音楽AIの内部表現から和音や調の構造を見つける
From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
この論文をやさしく読む
ひとことで言うと
曲を移調したときに音楽AIの内部の特徴がどう変わるかを調べ、和音や調に対応する特徴のまとまりを見つける研究です。
何に役立つ?
音楽モデルが学んだ表現を、音楽理論の概念と対応付けて理解するのに役立ちます。少数の基準例から概念の一群を解釈する手段として使うことが考えられます。
この研究の面白いところ
特徴を一つずつ名付けるのではなく、移調しても保たれる関係から意味を探します。音楽そのものの規則性を、AIの内部を分析する手がかりにしています。
どこまで分かった?
要旨での実験対象は二つの音楽基盤モデルです。復元された構造は和音、調、旋律パターンに対応しますが、音楽的な概念すべての解釈や、モデルの因果的な判断過程まで説明したとは述べていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音楽の基盤モデルが内部で何を学習したのかを、どのように理解できるだろうか。プロービングやスパースオートエンコーダ(SAE)など、解釈可能性に関する多くの手法は、構造に関する仮定を最小限にして、個々の特徴を特定することに重点を置いている。本研究では、多くの概念は孤立した特徴としてではなく、構造を持った関係として捉える方がよく理解できると論じる。この点は、音高と時間の空間の中で調性的な構造が組み立てられる音楽で特に顕著である。例えば、和音や調といった概念は、個別の特徴ではなく、構造化された集合として自然に表現される。和音の12通りの移調や、ある調の中の全音階体系がその例である。 本研究では特徴の特定から構造に基づく分析へと視点を移し、音楽基盤モデルが学習した内部表現に、複数の特徴にまたがる組織化された構造が現れるかを問う。そのために、移調を帰納バイアスとして用い、複数の見方から得たSAE表現の整列によって、順序を持つ軌道を導く枠組みを提案する。具体的には、音高をずらした入力のペアを生成し、そのSAE表現を整列させて、音高に関係する特徴の構造化された群を見つける。 実験の結果、この手法は二つの最先端の音楽基盤モデルにおいて、和音、調、旋律パターンに対応する軌道構造を復元することが示された。概念の一群全体を解釈するために必要な意味との対応付けは、少数の基準例など、ごくわずかで済む。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \textit{structured relations} rather than isolated features. This is especially prominent in music, where tonal structures are organized in the space of pitch and time. For example, concepts such as chords or keys are naturally expressed as structured sets (e.g., the 12 transpositions of a chord or the diatonic system within a key), rather than isolated features. In this study, \textbf{we shift from feature identification to structure-based analysis}, asking whether the learned inner representations of music foundation model emerge as organized structures over features. To this end, we introduce a framework that uses pitch transposition as an inductive bias to induce ordered orbits via multi-view SAE alignment. Concretely, we generate pitch-shifted input pairs and align their SAE representations to discover structured groups of pitch-related features. Experimental results show that this approach recovers orbit structures corresponding to chords, keys, and melodic patterns across two state-of-the-art music foundation models, while requiring only minimal grounding (e.g., a few anchor examples) to interpret entire concept families.
arXiv ID: 2610.01864 / 要約の誤りについて