arXiv論文メモ
新着一覧
cs.DB · 査読状況未確認

半構造化データから知識グラフを作る手法の評価基盤

Benchmarking Automated Knowledge Graph Construction from Semi-Structured Data

Tarek Al Mustafa and Birgitta König-Ries

この論文をやさしく読む

ひとことで言うと

半構造化データから作る知識グラフの品質を、複数の観点で測るベンチマークを作った。

何に役立つ?

知識グラフ構築手法を比べ、下流の質問応答に使えるか確認する際に役立つ。

この研究の面白いところ

構文や意味だけでなく、実際の質問に答えられるかまで一つの評価手順に含めた。

どこまで分かった?

7分野の10データセットと二つの参照システムで例示した。ほかの分野やシステムでの性能は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

知識グラフは従来の知識表現から、大規模言語モデルの下流課題を支える記憶まで、さまざまな用途で重要性を増している。ただし構築には多くの人手がかかり、自動化の手法が数多く提案されてきた。テキストを入力とする手法に比べ、半構造化データを扱う手法は少ない。そのため、対応付けの予測や生成された知識グラフの品質を判定する包括的なベンチマークと評価環境がない。知識グラフの品質は下流の用途に直接影響するため、構築手法をしっかり評価する仕組みが必要である。 本研究は、半構造化データからの知識グラフ構築に向けたベンチマークと評価手順を提示する。構文上の妥当性、意味的な正確さ、一貫性、簡潔さ、完全性に加え、能力確認用の質問に答えられるかという実用上の品質を組み合わせる。現実的な課題定義を与え、既存の評価枠組みを拡張し、対応付けを予測するシステムとRDFデータを生成するシステムの双方を評価できるようにする。また構築結果と下流での利用の評価を結び付ける。包括的な評価指標群と、7分野にまたがる専門家が整備した10のデータセットを提供し、二つの参照システムを用いて評価例を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Knowledge Graphs (KGs) play an increasingly important role in numerous applications ranging from traditional knowledge representation to serving as memory for LLMs to support downstream tasks. However, their construction is labor-intensive; thus, in recent years, numerous approaches for automizing this process have been proposed. Compared to the popularity of construction approaches that focus on textual input data, methods for semi-structured inputs remain underrepresented and as a result, no comprehensive benchmark and evaluation suite exists to judge the quality of mapping predictions and generated KGs. This is problematic, as a KG's quality has direct influence on the downstream applications it supports and thus, strong evaluation mechanisms for their construction are urgently needed. In this work, we thus focus on the evaluation of KG construction from semi-structured data and present a benchmark and evaluation pipeline for KG construction that combines the quality dimensions (1) syntactic validity, (2) semantic accuracy, (3) consistency, (4) conciseness, (5) completeness, and (6) pragmatic quality measured on a KG's ability to provide answers to competency questions. This work contributes a realistic task definition, extends current state of the art evaluation frameworks, allows evaluation of systems that predict mappings and RDF data alike, and combines evaluation of both KG construction and downstream usage. We provide a comprehensive metrics suite, provide ten expert-curated datasets from seven domains, and showcase evaluation using two reference systems.

著者のコメント

preprint - work in process document

arXiv ID: 2609.26985 / 要約の誤りについて