機能語による語彙の抽象化をTransformerは共有できるか
Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words
この論文をやさしく読む
ひとことで言うと
「研究者たちが論文を書いた」と「彼らがそれを書いた」のような文の共通構造を、Transformerの表現が捉えるか調べています。
何に役立つ?
言語モデルが名詞を代名詞へ置き換えた際にも文法や意味を一般化できるか、内部表現から評価する材料になります。
この研究の面白いところ
機能語は名詞より表現空間の中央に位置する一方、二種類の文は異なる部分空間にありました。片方だけの訓練では共通構造が現れず、両方を混ぜると現れる結果を示しています。
どこまで分かった?
事前学習済みTransformerの埋め込みと、構造情報を抽出する実験に基づく結果です。人間と同じ理解機構を持つと示したものではなく、要旨にはモデル別・言語別の広い検証範囲は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
代名詞、副詞などの機能語、例えばthey、her、somewhere、thereは、性や文法上の数といった性質だけでその文脈に十分な情報が得られる場合に、具体的な名詞や句を置き換えるためによく使われる。事前学習済みTransformerモデルは、人間のように使える形でこうした機能語を符号化しているのだろうか。言語モデルは、こうした語彙の抽象化に依存する「The researchers wrote the paper(研究者たちはその論文を書いた)」と「They wrote it(彼らはそれを書いた)」のような文の、統語的・意味的な対応を認識できるのだろうか。 これらの言語学的な問いを事前学習済みTransformerモデルの埋め込み空間へ写し、名詞の表現と、その名詞を置き換えられる代名詞・副詞の表現を比較する。単独の場合と、具体的な語彙を使う文および対応する機能語を使う文の中で比較する。さらに、これらの対応する文の埋め込みに、共通の統語・意味構造があるかを調べる。 機能語は名詞に比べて中心に位置するが、名詞とは区別されていることが分かった。これは、多様な文脈で代わりとなる語としての振る舞いと整合する。一方、具体語を使う文と機能語を使う対応文の埋め込みは、埋め込み空間の異なる部分空間に位置する。文の構造情報を抽出する実験では、いずれか一方の種類のデータだけで学習しても、共有構造は現れなかった。機能語データでは語彙が一貫しすぎ、具体語を使うデータでは語彙が多様すぎるためである。しかし、機能語文と具体語文を混ぜて学習すると、共有構造が現れる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Pronouns, adverbs and other functional words (such as they, her, somewhere, there) are often used in language to replace concrete nouns or phrases, when their properties - such as gender, grammatical number - provide sufficient information for the given context. Do pretrained transformer models encode such functional words in a manner that allows them to be used like humans do? Can language models recognize the syntactic and semantic parallelism of sentences such as "The researchers wrote the paper" and "They wrote it", which relies on such lexical abstraction? We map these linguistic questions into the embedding space of a pretrained transformer model, and compare representations of nouns, with the representations of the pronouns and adverbs that can replace these nouns, in isolation and in parallel lexicalized and functional sentences. We then probe for shared syntactic and semantic structure in the embeddings of parallel lexicalized and functional sentences. We find that functional words are located centrally compared to nouns, but are also distinct, which is congruent with their behaviour as place-holders in a wide variety of contexts. The analysis of the embeddings of parallel (lexicalized and functional) sentences show them inhabiting different subspaces of the embedding space. Experiments that distil the structural information of the sentence show that training on either type of data does not reveal the shared structure - because of the over-consistency of the vocabulary (in case of the functional data), and the too much variety (in case of the lexicalized versions). However, training with a mix of functional and lexicalized sentences, the shared structure emerges.
著者のコメント
16 pages, 11 figures
arXiv ID: 2609.19887 / 要約の誤りについて