多言語の化学特許を使った言語横断検索の評価
ChemCLIR-Bench: Benchmarking Cross-Lingual Information Retrieval in Multilingual Chemical Patents
この論文をやさしく読む
ひとことで言うと
化学特許を五言語で検索し、異なる言語をまたぐと関連文書が見つかりにくくなる程度を測る。
何に役立つ?
多言語の特許検索で埋め込みモデルを選ぶ際の評価材料になる。
この研究の面白いところ
最良モデルでもRecall@10が単一言語の0.72から異言語の0.53へ下がる。
どこまで分かった?
結果は構築した化学特許データセットと評価した8モデルに基づく。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
多国籍の産業では、重要な技術上の証拠が検索文とは異なる言語で書かれている場合があり、言語をまたぐ情報検索の重要性が増している。しかし従来の評価用データは、分野特有の言語横断検索や、集約した再現率では隠れる検索順位の深さと再発見可能性の失敗を十分に捉えていない。本研究は、特許データを中心に化学分野の言語横断検索を評価する。 Google Patentsと欧州特許庁(EPO)のデータから、東西の主要言語を含む五言語の多言語データセットを構築し、実際の産業文書の多様さと複雑さを反映させる。このデータで、最新の埋め込みモデル8種類を体系的に評価した。最も成績のよいモデルでも、上位10件の再現率Recall@10は、単一言語での0.72から異言語での0.53へ下がった。関連文書の順位が異言語検索で下がるなど、検索の深さも大きく悪化する。単一言語では高性能な一部の多言語埋め込みモデルも、質問と文書の言語が異なると性能が急落した。これらは、分野に特化した言語横断検索における現行手法の限界を示し、モデル選択と評価のための統制された診断枠組みを提供する。データとコードはhttps://github.com/MohammadKhodadad/Multi-Lingual-QACで公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Cross-lingual information retrieval (CLIR) is increasingly important in multi-national industries, where critical technical evidence may exist in a different language than the query. However, existing benchmarks do not adequately capture domain-specific cross-lingual retrieval or the retrieval-depth and recoverability failures that aggregate recall hides. In this work, we benchmark CLIR in the chemical domain, with a focus on patent data. We construct a multilingual dataset from Google Patents and the European Patent Office (EPO) data, spanning five languages (covering major Eastern and Western languages) and reflecting the diversity and complexity of real-world industrial documentation. Using this dataset, we systematically evaluate eight state-of-the-art embedding models for cross-lingual retrieval. Our results show a substantial performance gap between monolingual and cross-lingual settings: for the best-performing model, Recall@10 drops from 0.72 to 0.53 in cross-lingual setting. Retrieval depth also degrades significantly, with relevant documents ranked lower across languages in cross-lingual scenarios. Furthermore, some multilingual embedding models that perform strongly in monolingual settings exhibit sharp declines when queries and documents are in different languages, providing practical insights for model selection in cross-lingual use cases. These findings highlight critical limitations of current approaches and emphasize the need for more robust cross-lingual retrieval methods in domain-specific settings. Our benchmark provides actionable insights for model selection and establishes a controlled diagnostic evaluation framework for CLIR over industrial technical text. Data and code are publicly available at https://github.com/MohammadKhodadad/Multi-Lingual-QAC.
著者のコメント
Accepted to the EMNLP 2026 Industry Track. 22 pages including references and appendices; 7 pages of main text
arXiv ID: 2609.23231 / 要約の誤りについて