文化・言語・地域の注釈で学習文書を監査するデータセット
FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing
この論文をやさしく読む
ひとことで言うと
309億件のウェブ文書と277件の評価課題に共通の文化・言語・地域の注釈を付けた。
何に役立つ?
文化的な現象が学習データにどれだけ含まれ、評価課題でどれだけ試されるかを比べる監査に使える。
この研究の面白いところ
文書とベンチマークを同じ分類体系で結び、URLから地域を付けた文書は79.2億件に達した。
どこまで分かった?
全309億文書に注釈を付けたが、空でない地域ラベルを割り当てられたのは25.61%だった。要旨はこのデータセットを使ったモデル性能の改善を示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデルの文化に関する評価範囲と頑健性は、事前学習の文書群と文化に関するベンチマークが、比較できる共通のメタデータで整理されていないため診断しにくい。ベンチマークは言語、地域、その土地特有の慣行に根ざした現象を対象にすることが増えているが、ウェブ規模の文書群は通常、言語だけで分類される。文化・言語・地域を共通軸にすれば、対象とする文化現象が事前学習データに含まれるか、ベンチマークで評価されるか、あるいはその両方かを監査できる。そこでFineWebとFineWeb-2から派生した大規模注釈付きデータセットFineWeb-CLaRを導入し、ウェブ文書を共通の文化・言語・地域軸に配置して、文書群の監査とベンチマークとの対応付けを可能にする。FineWebとFineWeb-2の全309億文書に、URLから推定した地域ラベルと文化的話題の出所情報を付与した。地域を判定する仕組みは、文書の25.61%に当たる79.2億件へ空でない地域を割り当てた。文化的話題の分析では、地域固有の話題を導き、Liuら(2025)の文化分類体系の末端14分類に対応付け、文書群側で比較できる地域別話題分布(LTD)を作った。また、文化に関する自然言語処理ベンチマーク277件にも、同じ分類体系、対象言語、対象地域の注釈を付けた。これらを合わせることで、事前学習文書群側に存在する根拠と、ベンチマーク側の評価範囲を直接比較できる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while web-scale corpora are usually organized only by language. A shared culture-language-region layer makes these resources comparable, enabling audits of whether a target cultural phenomenon is represented in pretraining data, evaluated by benchmarks or both. To this end, we introduce FineWeb-CLaR, a large-scale annotated dataset derived from FineWeb and FineWeb-2 that places web documents on a shared culture-language-region axis for corpus auditing and benchmark alignment. FineWeb-CLaR annotates the full 30.9B-document collection from FineWeb and FineWeb-2 with URL-derived region labels and cultural-topic provenance. Our region resolver assigns a non-empty region to 25.61% of documents (7.92B). For cultural-topic analysis, we induce locale-specific topics and project them onto the 14 leaves of the Cultural Taxonomy of Liu et al. (2025), producing Locale Topic Distributions (LTDs) for corpus-side comparison. We also annotate 277 cultural NLP benchmarks with the same taxonomy, language coverage, and region coverage. Together, these resources enable direct comparison between corpus-side pretraining evidence and benchmark-side evaluation coverage.
著者のコメント
accepted to EMNLP 2026 (Main)
arXiv ID: 2609.25298 / 要約の誤りについて