社会集団に関する投稿5億2700万件の公開コーパスISAAC
The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale
この論文をやさしく読む
ひとことで言うと
社会集団に関するReddit投稿を長期間・大規模に集め、地域や意味ラベルを付けた公開データ基盤。
何に役立つ?
社会集団への言及が時間や地域でどう変わるかを比較する実証研究に役立つ。
この研究の面白いところ
5億2700万件以上を収め、無関係な投稿を全体と各区分で10%未満に抑える選別を行った。
どこまで分かった?
対象は2007~2023年の英語Reddit投稿。推定居住地域や自動ラベルには推定が含まれ、他の媒体への適用結果は要旨で示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Illinois Social Attitudes Aggregate Corpus(ISAAC)を紹介する。これは、人種、性的指向、年齢、障害、体重、肌の色という六つの社会集団上の区別に関係する英語のReddit投稿を5億2700万件以上収録した、公開・拡張可能で利用しやすいコーパスである。対象期間は2007年から2023年までの17年間。複数段階で人が監査する選別工程により、関係のない内容が整備済みデータに占める割合を、全体でも各区分でも10%未満に抑えた。各投稿には、利用者の推定居住地域と、道徳的評価、感情の極性、情動、言語的な一般化を含む、既存および独自の検証済みの意味ラベルをアルゴリズムで付与した。 ISAACと、オンライン検索行動、全国・地域での大きな社会的出来事に伴う時期別の急増、長期的な世論の変化といった社会全体の動向が対応する複数の証拠によって、コーパスの妥当性を確かめた。統一された公開基盤により、研究の分断を減らし、大規模な実証研究の再現を容易にする。具体的には、集団区分間の比較、集団についての言説が長期的にどう変わるかの精密な追跡、地域差と地域ごとの世論・政策結果との対応付けができる。公開された拡張可能な処理工程により、他のプラットフォーム、言語、社会的区分への展開も容易になる。利用方法は、操作型のウェブサイトとラベル付け用ウェブアプリに加え、SQLの試用環境、Pythonパッケージ、Hugging Faceにも対応する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We introduce the Illinois Social Attitudes Aggregate Corpus (ISAAC), an open, modular, and accessible corpus of 527 million+ English-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17-year period from 2007 to 2023. A multi-step, human-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off-the-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro-level societal trends, such as online search behavior, temporal spikes during major societal events (both nationally and regionally), and long-term shifts in public attitudes. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale. Specifically, ISAAC allows investigators to perform cross-category comparisons, conduct high-precision tracking of long-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories. To accommodate various research needs, ISAAC is accessible both without coding through a point-and-click website and labeler web-apps, and programmatically via an SQL playground, a Python package, and HuggingFace.
著者のコメント
Submitted to Behavior Research Methods
arXiv ID: 2609.27059 / 要約の誤りについて