親族呼称を選べても生成できない言語モデル
Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms
この論文をやさしく読む
ひとことで言うと
親族呼称を選択肢から選ぶ成績は高くても、自分で正しい語を生成する成績は低いと報告した。
何に役立つ?
多言語モデルを親族関係の実際の表現に使う場合の評価方法を考えるのに役立つ。
この研究の面白いところ
同じ親族関係と言語の組で選択と生成を比べ、言語ごとの父系優位の違いも確認した。
どこまで分かった?
選択と生成の差は評価形式の違いを示すが、語彙知識が内部に完全に保たれている証拠ではない。対象は三言語、五モデルである。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
これまでの研究では、多言語の親族関係を大規模言語モデル(LLM)が理解しているかを、多肢選択のベンチマークで測り、認識の問題として扱ってきた。本研究ではそれに代わり、五つのオープンウェイトLLMに対し、ヒンディー語、タミル語、韓国語の三つの非西洋言語で、二種類のコミュニケーション課題における親族呼称を生成させ、対応する選択肢付きの選択課題を基準として比べる。同じ親族関係と言語の組み合わせでは、GPT-OSS 120Bは有効な75件のうち90.67%で正しい語を選択したが、対応する生成試行で受け入れられる語を出せたのは36.00%だった。Llama 3.3 70Bにも同じ傾向が見られ、選択77.92%に対し生成24.24%だった。 四つの選択肢がある条件では候補語が表示され、文字表記を自力で生成する必要もないため、この差は評価形式による隔たりと解釈される。語彙知識が保たれていることの直接の証明ではない。関係を明示したL3プロンプトでは、正答率はGLM-5.1の72.29%からLlama 3.3 70Bの24.24%まで大きく異なった。父系に関する優位性は言語に依存し、ヒンディー語では大きいが韓国語では弱いか逆転し、タミル語で同じ語を共有する親族関係の組は測定上の変動を確認する対照となる。文化に固有の親族呼称の生成は、関係が明示されてもなお難しく、多肢選択と併せて生成による評価も行う必要性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weight LLMs to generate kinship terms in three non Western languages (Hindi, Tamil, and Korean) across two communicative tasks and pair this with a matched option-supported selection baseline. On identical relation language cells, GPT OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in 36.00% of the corresponding attempts; Llama 3.370B shows the same pattern (77.92% versus 24.24%). Since the four-option condition displays the candidate terms and does not require script production, the difference is interpreted as an evaluation format gap rather than direct proof that lexical knowledge is intact. On explicitly specified L3 prompts, accuracy varies sharply, from GLM-5.1 at 72.29% to Llama-3.370B at 24.24%. The paternal-lineage advantage is language specific; it is large in Hindi but weak or reversed in Korean, while Tamil shared-term pairs provide a control for measurement variation. These results show that culturally specific kinship generation remains difficult even when the relationship is explicitly stated and motivate generation-based evaluation alongside multiple-choice testing.
著者のコメント
Accepted at (ORACLE Workshop), EMNLP 2026
arXiv ID: 2609.26942 / 要約の誤りについて