arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

韓国語の新語を言語モデルが理解できるか測る

KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

Soha Lee, Soojin Lee, Heesung Yang, Hyunju Song, Hyunji Lee, Jinsan An, Jeongwan Shin, Jin Hyun Park, Jun Lee, Hyeyoung Park, Kilim Nam

この論文をやさしく読む

ひとことで言うと

最近の韓国語の新語を、言語モデルが意味や成り立ちまで理解できるか調べる評価データセットです。

何に役立つ?

既存語中心の試験では分かりにくい、新しい語彙への対応力を測る用途があります。人間の基準と比較し、語の構成要素や定義のどこでモデルがつまずくかを評価できます。

この研究の面白いところ

2020年以降のオンラインニュースで確認された1,785語を専門家が点検し、用例、語形成の分析、辞書形式の定義を付けています。韓国語で語と機能的な形態素が組み合わさる性質を評価に取り込んでいます。

どこまで分かった?

四つの課題で、語源的な構成要素の復元、意味分類、正確な定義生成に限界が報告されています。要旨にはモデル別の得点や、韓国語能力全般への一般化を裏付ける結果は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自然言語では新しい語や意味が次々に生まれるにもかかわらず、大規模言語モデル(LLM)は一般に固定的なベンチマークで評価されている。既存の韓国語ベンチマークは定着した語彙が中心であるため、近年の語彙変化を十分に網羅していない。また、英語を基準にした設計のため、内容語と機能形態素が生産的に結び付くという韓国語の言語類型上の特性を評価しにくい。 本論文では、LLMによる韓国語新語の理解を評価するベンチマークKoNeoBenchを導入する。KoNeoBenchは、2020年以降のオンラインニュースで用例が確認された韓国語の新語1,785語を基に、辞書編纂の専門家による精査を経て構築した。各項目には、用例、語形成の分析、辞書形式の定義を収録する。この資源を基に四つのタスクを定義し、最近のモデルの結果を人間の基準成績とともに報告する。 実験から、現行のLLMには、語を構成する元の要素の復元、意味カテゴリーの区別、正確な定義の生成に明確な限界があることが分かる。これらの結果は、近年の韓国語の語彙変化のうち、現在のLLMにとって依然として難しい具体的な側面を明らかにする。KoNeoBenchは https://github.com/bcmilab/ko-neobench/ で公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary and therefore provide limited coverage of such recent lexical change, and their English-oriented design makes it difficult to assess the typological properties of Korean, in which content words combine productively with functional morphemes. In this paper, we introduce KoNeoBench, a benchmark for evaluating LLMs' understanding of Korean neologisms. KoNeoBench is built on 1,785 Korean neologisms attested in online news since 2020 and curated through expert lexicographic review. Each entry provides usage examples, word-formation analyses, and dictionary-style definitions. Based on this resource, we define four tasks and report results on recent models, together with a human baseline. Our experiments show that current LLMs exhibit clear limitations in recovering source components, distinguishing semantic categories, and generating accurate definitions. These results reveal specific aspects of recent Korean lexical change that remain challenging for current LLMs. KoNeoBench is available at https://github.com/bcmilab/ko-neobench/ .

著者のコメント

Accepted to Findings of EMNLP 2026. Code and data are available at the project repository

arXiv ID: 2609.19916 / 要約の誤りについて