ヒンディー語の抽出型要約を再現し、評価方法を検証
Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study
この論文をやさしく読む
ひとことで言うと
ヒンディー語の抽出型要約を再現したところ、2つの評価データでは記事の冒頭3文を使う単純な方法を上回れなかった。
何に役立つ?
ヒンディー語の要約システムを比較する際、評価データが冒頭文の再利用を強く反映していないか確認する手掛かりになる。
この研究の面白いところ
言語に合わせた処理と評価器の照合に加え、特徴量の除去実験と文の選択傾向から、性能差の理由を調べた。
どこまで分かった?
結論は調べた2つのコーパスと抽出型手法の比較に基づく。冒頭以外を適切に選ぶ要約が実際に役立たないことを示したわけではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Mohd、Jan、Shah(2020年)の分布意味論に基づく抽出型要約手法を再現し、言語固有の各処理をデーヴァナーガリー文字に適したものへ置き換えてヒンディー語に適応した。XL-Sumのヒンディー語部分とFIRE ILSUM 2.0 Hindiという独立した2つのコーパスで評価した。評価には、XL-Sum著者の多言語スコアラーと照合済みの、デーヴァナーガリー文字に対応するROUGEを使い、比較にはすべて1,000回の再標本化による対応ありブートストラップ法を用いた。 公表された等重み設定の再現システムは、両コーパスで冒頭3文を使う基準手法より有意に悪く、ROUGE-1 F値はXL-Sumで0.042、ILSUMで0.265下回った。特徴量を一つずつ除く分析では、寄与したのは文の位置だけだった。位置だけを使うと冒頭文の基準手法と完全に一致し、位置を除くと最も弱い設定になった。検証データで重みを調整しても、冒頭3文と同等になるのが最善で、上回ることはなかった。TextRankも同様に失敗し、この結果は一つの実装に限らず手法群に関わるものと分かった。 選択文の分析では、残る特徴量が長く固有表現の多い本文の文を選ぶ一方、参照要約は記事の冒頭を再利用していた。現在のヒンディー語のベンチマークでは冒頭以外の内容選択が評価されにくいため、その目的に合った評価資料が必要だと論じる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We replicate the distributional-semantics extractive summarisation method of Mohd, Jan and Shah (2020) and adapt it to Hindi, substituting a Devanagari-appropriate component at every language-specific step. The system is evaluated on two independent corpora --- the Hindi portion of XL-Sum and FIRE ILSUM 2.0 Hindi --- under a Devanagari-aware ROUGE implementation validated against the XL-Sum authors' own multilingual scorer, with all comparisons drawn as 1000-resample paired bootstraps. In its published equal-weight configuration the replicated system is significantly worse than a three-sentence lead baseline on both corpora, trailing Lead-3 by 0.042 ROUGE-1 Fon XL-Sum and by 0.265 on ILSUM. A feature ablation shows that sentenceposition is the only feature that contributes: position alone reproduces the lead baseline exactly, removing position gives the weakest configuration,and a validation-tuned weighting can at best equal Lead-3 and never exceed it. TextRank fails identically, making this a class-level rather than an implementation-level result. A selection analysis shows the remaining features steer extraction towards long, entity-dense body sentences while the references reuse the article lead.Current Hindi benchmarks therefore cannot reward non-lead content selection, motivating purpose-built evaluation resources.
arXiv ID: 2609.29090 / 要約の誤りについて