arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

複合語の意味を近い語群からの類推で予測する

Compound interpretation is based on analogy

Tian Shen and Harald Baayen

この論文をやさしく読む

ひとことで言うと

複合語の意味を、同じ構成要素を持つ別の複合語との類似関係から計算するモデルを調べています。

何に役立つ?

語の意味の予測だけでなく、人が複合語を理解する仕組みについて、異なる理論を比較するために役立ちます。語彙判断の反応時間との対応も評価しています。

この研究の面白いところ

学習するパラメータを持たない局所的な類推モデルが、学習済みの大域的な変換を使う比較モデルを多くの条件で上回った点が特徴です。

どこまで分かった?

評価対象は標準中国語の複合語です。3文字複合語では優位性が得られず、複合語族の小ささと大きさの不均衡が制約として挙げられています。反応時間の予測結果だけで、人の処理機構を一意に確定したわけではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

複合語の意味を構成要素の意味からどのように予測するのが最もよいかは、語彙意味論の計算モデルにおける中心的な問いであり続けている。異なる計算モデルの比較は、複合語を理解する際に意味情報がどのように組み合わされるかについて、複数の説明を評価する手段となる。 本研究では、複合語類推モデル(Compound Analogy Model:CAM)という新しいモデルを提案する。このモデルは、複合語を構成する各要素の埋め込みに、それぞれの要素を含む複合語族の平均シフトベクトルを加えることで、複合語の埋め込みを予測する。得られたモデルはパラメータを持たず、意味空間の局所的な類推構造を利用する。 標準中国語の複合語について、CAMをCAOSSモデルと比較評価した。CAMは、学習データと評価用に取り置いたデータの両方で、一貫してCAOSSより高い予測精度を達成した。ただし3文字の複合語は例外であり、各構成要素の複合語族が小さいことと、2つの構成要素の複合語族の大きさに著しい偏りがあることの両方によって、類推による汎化が制約される。既知の複合語から新しい複合語への汎化をよりよく近似する、頻度に基づく学習・テスト分割で評価しても、CAMの優位性は維持された。 両モデルの認知的な妥当性を評価するため、モデルから導いた意味的指標が、2文字複合語の視覚的語彙判断における反応時間を予測するかをさらに調べた。CAM由来の予測変数は、CAOSSモデル由来の予測変数よりも反応時間をよく予測した。これらの知見は、複合語の意味が、学習された大域的な線形変換の適用よりも、局所的な類推による汎化としてよく特徴付けられることを示し、意味の類推構造が、複合語理解の認知的に妥当な基盤になることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

How compound meanings are best predicted from constituent meanings remains a central question in computational models of lexical semantics. Comparing different computational models provides a way to evaluate alternative accounts of how semantic information is combined during compound comprehension. We propose a new model, the Compound Analogy Model (CAM), that predicts a compound's embedding by adding its constituent embeddings together with the average shift vectors of the two constituents' compound families. The resulting model is parameter-free and exploits local analogical structure in the semantic space. We evaluated CAM against the CAOSS model on Mandarin Chinese compounds. CAM consistently achieved higher prediction accuracy than CAOSS on both training and held-out data, with the exception of three-character compounds, for which analogical generalization is constrained by both small constituent families and a pronounced imbalance in family size between the two constituents. The advantage of CAM remained when evaluation was based on frequency-defined train-test splits that better approximate generalization from familiar to novel compounds. To assess the cognitive plausibility of the two models, we further examined whether model-derived semantic measures predict visual lexical decision latencies for two-character compounds. Predictors derived from CAM provided improved prediction for response latencies compared to predictors derived from the CAOSS model. These findings indicate that compound meaning is better characterized as local analogical generalization than as the application of a learned global linear transformation, and demonstrate that analogical semantic structure provides a cognitively plausible basis for compound comprehension.

arXiv ID: 2610.01688 / 要約の誤りについて