arXiv論文メモ
新着一覧
cs.IR · 査読状況未確認

投稿文から項目を見つけて構造化データを作る

Learning to structure data from user-generated thematic corpora

Elishay Avram, Oren Glickman, Elad Yom-Tov

この論文をやさしく読む

ひとことで言うと

投稿を読む前に抽出項目を決めなくても、LLMが項目の発見、整理、値の抽出まで行う方法です。

何に役立つ?

大量のテーマ別投稿を、属性ごとの表として分析しやすくする用途が考えられます。評価では健康関連のReddit投稿を使っています。

この研究の面白いところ

項目を見つける能力と、型を付ける能力、値を抜き出す能力を別々に評価しています。属性の一致率は人同士の一致率に近く、小型モデルへの切り替え時の精度と費用も検討しています。

どこまで分かった?

人との比較と、大型LLMを基準にした比較は異なる評価です。61%の属性一致、82%の型精度、F1 0.8はそれぞれ別の指標で、すべての情報を完全に構造化できたことを意味しません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ソーシャルメディアのコミュニティなど、特定テーマに関するコーパスには、構造化できるデータを記述した非構造化テキストが含まれる。例えば、ソーシャルメディアのデータで言及される個人の属性、行動、経験などである。関係する属性は暗黙的で、分野に依存し、事前には分からないことが多いため、構造化データの抽出は難しい。本研究では、事前定義されたオントロジーなしに、分野固有の属性スキーマを発見・抽出する、完全自動の反復的な枠組みを提案する。大規模言語モデル(LLM)を用いて候補属性を導出し、意味の重なる属性を順に統合して、構造上の型を割り当てる。これによりオントロジーを作り、コーパスから得た値で埋めることができる。また、大きなLLMと比べた精度低下を推定しながら、小さなLLMを値の抽出に利用できる。 健康に関するRedditの5コミュニティで評価した。発見された属性と人が特定した属性の一致率は61%で、独立した注釈者間の一致率62%に近かった。ほとんどの場合、アルゴリズムは10回未満の反復で安定した属性集合に収束する。構造上の型の割り当て精度は82%、値の抽出は人の注釈との比較でF1スコア0.8に達した。4系列のLLMにわたり、高能力の参照LLMを基準に評価した小型の指示調整モデルでは、モデル規模の増加に伴い抽出性能が統計的に有意に改善し、精度とコストの適切な折り合いを付ける判断を支える。結果は、人が特定するものに匹敵する属性を自動で発見でき、高品質な構造化データセットを経済的かつ大規模に作成できることを示す。事前定義のオントロジーを不要にする反復的なモデル駆動型スキーマ導出は、テーマ別コーパスを分析するための実用的で拡張可能な基盤を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Thematic corpora, such as social media communities, contain unstructured text describing data that could be made structured. These include, for example, personal attributes, behaviors, and experiences mentioned in social media data. Extracting structured data is challenging as relevant attributes are often implicit, domain-dependent, and unknown in advance. We propose a fully automated, iterative framework for discovering and extracting domain-specific attribute schemas without a predefined ontology. Using large language models (LLMs), the framework induces candidate attributes, sequentially consolidates semantically overlapping attributes, and assigns a structural type. These enable creating an ontology and populating it with values from the corpus. The framework also enables the use of smaller LLMs for value extraction with estimable accuracy loss compared to large LLMs. We evaluate the framework on 5 health-related Reddit communities. Discovered attributes achieved 61% agreement with human-identified attributes, close to the 62% agreement between independent annotators. In most cases, the algorithm converges to a stable attribute set in fewer than 10 iterations. Structural type assignment achieves 82% accuracy, and value extraction reaches an F1 score of 0.8 compared to human annotations. Across four LLM families, smaller instruction-tuned models show statistically significant improvements in extraction performance with model scale when evaluated against a high-capacity reference LLM, supporting informed accuracy-cost trade-offs. These results show that attributes comparable to those identified by humans can be discovered automatically, enabling the creation of high-quality structured datasets economically and at scale. By removing the need for predefined ontologies, iterative model-driven schema induction offers a practical and scalable foundation for mining thematic corpora.

arXiv ID: 2610.01463 / 要約の誤りについて