音声言語モデルは発音の特徴を音と文字で同じように表すか
Do Audio Language Models Hear and Read Distinctive Features Alike?
この論文をやさしく読む
ひとことで言うと
音声と言語を扱うモデルが、発音上の違いを聞く場合と文字で読む場合に同じように内部表現するか調べた研究。
何に役立つ?
音声言語モデルの内部表現を比較し、音声入力と文字入力の対応を評価する方法として役立つ。
この研究の面白いところ
無作為な音素対を基準にすることで、もともと存在する系列間の一致を差し引いて評価した。補正後に基準を超えた特徴は二つのQwen2.5-Omniモデルの有声性のみだった。
どこまで分かった?
分析対象は6モデル、7特徴、15言語であり、すべてのモデルや発音特徴に一般化した結果ではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声言語モデルは音声と文字を一つのデコーダーに通す。本研究では、音素を聞いたときと読んだときに、そのデコーダーが弁別的特徴を同じ方向に表現するかを問う。一つの特徴だけが異なる音素の最小対について、両者の平均表現の差を求める。その差を平均して音声と文字の各系列での方向を定め、二つの方向のコサインを測る。両系列は任意の音素対について既にある程度一致するため、測定値をゼロではなく、無作為に組み合わせた対から作る基準と比較する。 11語族の15言語、7種類の特徴、6モデルに適用した。多重検定の補正後に基準を上回ったのは、二つのQwen2.5-Omniモデルにおける有声性だけだった。基準値はモデル間で7倍異なった。6モデル中3モデルでは、測定に十分な最小対がある14言語にわたり、音声内の有声性が一つの方向を持った。そのうち2モデルでは、どの言語対でも方向が一致した。どの系列で特徴が表現されるかを予測したのは、モデルの大きさではなくモデル系列だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members' mean representations. Averaging those offsets gives a direction for each stream, and we measure the cosine between the two. Because the two streams already agree about arbitrary phoneme pairs, we compare every measure against a reference built from random pairings rather than against zero. We apply this to 6 models, 7 features and 15 languages from 11 families. Only voicing in the two Qwen2.5-Omni models exceeds that reference after correction for multiple testing, and the reference varies by a factor of seven between models. In three of the six models, voicing has one direction in audio across the 14 languages with enough minimal pairs to measure it, and every language pair agrees in two of them. The model family, not the model size, predicts which stream represents a feature.
arXiv ID: 2609.30167 / 要約の誤りについて