arXiv論文メモ
新着一覧
cs.CL / cs.DL · 査読状況未確認

構造類似性を用いた論文の文順分類の多言語転移改善

Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity

Kazuhiro Yamauchi, Marie Katsurai

この論文をやさしく読む

ひとことで言うと

論文の文を目的・方法などへ分類するモデルを別言語へ移すとき、言語の近さより文章構成の似方が役立つかを調べています。

何に役立つ?

英語以外の訓練データが少ない科学文献を構造化し、多言語の文献利用を支える方法です。

この研究の面白いところ

5つの学術データベースから13の非英語言語のデータを作成しました。言語的距離に一貫した予測力はなく、ラベル分布や構成の類似性に弱いが一貫した関係が見られます。

どこまで分かった?

相関は強いとはされていません。提案法の最良の組合せは同領域で強いエンコーダーと同等、未見言語への転移で上回るという、評価設定ごとの違いがあります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文順分類(SSC)は学術出版物を構造化するために重要な課題であり、英語以外の言語へSSC研究を広げることは、多言語デジタル図書館での科学知識へのアクセスを改善しうる。非英語言語では学習データが不足するため、言語間転移は有望なアプローチである。他の自然言語処理課題での先行研究は、ソース言語と対象言語の言語的類似性を捉えることの利点を示してきた。しかしSSCは、言語の違いにかかわらず各言語で一貫して現れる、ラベル列や位置の規則性など談話レベルのパターンに本質的に依存する。SSCの転移成功を決める要因を調べるため、5つの学術データベースから収集した13の非英語言語を含む多言語SSCデータセットを構築した。 エンコーダベースモデルと生成モデルを用いた言語間転移実験では、言語的近さは転移性能を一貫して予測する力を持たない一方、修辞構成の構造的類似性はモデルをまたいで弱いが一貫した正の相関を持つことが分かった。ソース言語での性能を調整すると、ラベル分布の類似性が最も一貫した予測因子となった。この知見に基づき、生成モデルで構造情報を明示的に利用する3つの方法を提案する。言語内評価では最良の組み合わせが最も強いエンコーダベースラインと同等に達し、学習時に見ていない言語への転移では最も強いエンコーダベースラインを上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDFDOI

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.

著者のコメント

Accepted at JCDL 2026 (ACM/IEEE Joint Conference on Digital Libraries), Frisco, TX, USA, October 13-16, 2026. 12 pages, 5 figures, 9 tables. DOI: 10.1145/3805696.3846040

arXiv ID: 2609.19650 / 要約の誤りについて