arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

ウェブ上の質問頻度は人の知りたいことを表すのか

You Can Tell Who's Asking: What the Web's Questions Are Made Of, and Where They Come From

Calvin Zhou, Vincent McCloskey, Krishna Srinivasan

この論文をやさしく読む

ひとことで言うと

ウェブに同じ質問文が多くあるからといって、多くの人が実際にそれを尋ねたとは限らないことを、大量のページで調べています。

何に役立つ?

質問データを学習や検索評価、需要調査に使う際、定型文の重複や収集元の偏りを考慮する材料になります。

この研究の面白いところ

134億件は質問の種類数ではなく出現件数です。質問内容の種類より、長さや周辺文脈が由来の判別に効いたと報告しています。

どこまで分かった?

対象はクロール可能なウェブと FineWeb の収集範囲です。AUC は正解率ではなく、商業 FAQ との識別は0.554にとどまります。79%は割合の相対的な減少であり、79パーセントポイントの減少ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ウェブから収集した質問は、学術界と産業界の両方で、人々が何を知りたいかの代理指標として使われている。質問応答の学習データ、検索ベンチマーク、コンテンツ戦略では、ページ上の質問が人間の意図を反映すると仮定される。本研究では、2013〜2025年の FineWeb の110スナップショットから134億件の質問の出現を抽出して、この仮定を大規模に検証し、3つの知見を報告する。 第一に、誰が尋ねているかを見分けられる。質問の由来、すなわち掲載ホストやページは、質問の形式に信号を残す。ロジスティックモデルは、質問の種類よりも長さや周辺文脈によって、実際の利用者による質問と、テンプレート化または人工的に作られた質問を AUC 0.725 で区別できる。ただし、商業的な FAQ の文章との区別では0.554にとどまる。第二に、質問の頻度は需要を測らない。頻出上位1,000件の70%超は定型文またはテンプレート化された質問であり、出現数が測るのは、ある文字列が何回公開されたかであって、何回尋ねられたかではない。 第三に、12年間で、出現全体に占める実際の質問の割合は79%減少した。クロール対象の構成を調整すると減少は42〜56%であり、質問の長さと文脈も減少した。本研究は、ウェブ上の質問の由来を、出現単位で時間を追って測定する初の研究を提示し、クロール可能なウェブの質問が、人間が尋ねるものから、機械に読ませるために製造されるものへ移ってきたことを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Questions scraped from the web are used across academia and industry as a proxy for what people want to know. Across QA training data, retrieval benchmarks, and content strategy, questions on a page are assumed to reflect human intent. We test this assumption at scale by extracting 13.4B question occurrences across 110 FineWeb snapshots (2013-2025), and report three findings. First, you can tell who is asking: provenance (the host/page of questions) leaves a signal in question form, and a logistic model can separate genuine user questions from templated/manufactured ones at AUC 0.725 via length and surrounding context rather than question type, though only 0.554 against commerce FAQ writing. Second, question frequency does not measure demand: the most-frequent questions are boilerplate/templated (over 70% of the top thousand), so occurrence counts measure how often a string was published and not how often it was asked. Third, over twelve years the genuine share of occurrences fell by 79% (42-56% after controlling for crawl composition), with question length and context decreasing. We present the first diachronic, occurrence-level measurement of web question provenance, and find the crawlable web's questions have shifted from being asked by humans toward manufactured for machines to read.

著者のコメント

Accepted to the 13th Web as Corpus Workshop (WaC-13) at EMNLP 2026. 14 pages, 4 figures. Code and data: https://github.com/bodhiumlabs/tell-whos-asking

arXiv ID: 2609.24106 / 要約の誤りについて