arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

質問から作るオントロジーの語彙・構造・公理を分けて評価

CQ4OE: A benchmark for assessing LLM-assisted ontology generation from competency questions

Jiayi Li, Ziyuan Wang, Daniel Garijo and María Poveda-Villalón

この論文をやさしく読む

ひとことで言うと

質問に答えられる知識体系をAIに作らせる課題で、用語を拾う能力と、関係や論理を正しく構成する能力を分けて測ります。

何に役立つ?

オントロジー生成モデルを比較するとき、単語が合っているだけの出力を、論理構造も適切な出力と混同せず評価できます。

この研究の面白いところ

各質問と必要なクラス・関係・公理の対応を正解データに明示し、用語と全体構造の二段階で評価します。

どこまで分かった?

用語評価は99問、全体評価は118問で、対象が異なります。要旨には各モデルの具体的スコアや、すべての領域への一般化性能は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

能力質問(CQ)からのオントロジー生成は、オントロジー工学の中心的な工程だが、多くの労力を要する。大規模言語モデル(LLM)は有望な自動化能力を持つものの、現在の評価は断片的である。タスクの定式化は不統一で、正解標準にはCQごとの細かな来歴が不足しがちであり、評価指標は語彙の重なりと構造的・論理的な適切さを混同し、参照オントロジーも評価用CQを中心に明示的に設計されているとは限らない。これらの問題に対し、CQからのLLMによるオントロジー生成を系統的かつ再現可能に評価するベンチマークCQ4OEを提示する。各オントロジーについて、CQに基づく正解OWLオントロジーを構築し、それぞれのCQを、回答に必要なクラス、プロパティ、公理へ結び付ける来歴を明示する。この資源から、相補的な二つの評価課題を定義する。CQ2Termは99個のCQを対象に、CQ固有のクラスとプロパティの予測を用語レベルで評価する。CQ2Ontoは118個のCQを対象に、階層、プロパティのモデル化、公理レベルの構造を含むオントロジー全体を評価する。九つのLLMを用い、ゼロショット、反復、マルチエージェントの生成戦略でCQ4OEの実験を行う。その結果、LLMはオントロジーを構築するより、明示された語彙を取り出す方が安定しており、特にプロパティのモデル化、階層構築、公理生成に課題があることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Ontology generation from Competency Questions (CQs) is a central yet labor-intensive phase of Ontology Engineering. While large language models (LLMs) offer promising automation capabilities, current evaluations remain fragmented. Task formulations are heterogeneous, gold standards often lack fine-grained CQ provenance, metrics conflate lexical overlap with structural and logical adequacy, and reference ontologies are not always explicitly designed around the evaluation CQs. Here, we address these limitations with CQ4OE, a benchmark for the systematic and reproducible evaluation of LLM-based ontology generation from CQs. For each ontology in the benchmark, we build a CQ-driven gold OWL ontology with explicit provenance linking each CQ to the classes, properties, and axioms required to answer it. From this resource, we define two complementary evaluation tasks. CQ2Term supports term-level evaluation of CQ-specific class and property prediction over 99 CQs, and CQ2Onto supports ontology-level evaluation over 118 CQs, including hierarchy, property modeling, and axiom-level structure. We demonstrate CQ4OE with experiments using nine LLMs under zero-shot, iterative, and multi-agent generation strategies, showing that LLMs recover explicit vocabulary terms more reliably than creating ontologies, particularly in property modeling, hierarchy construction, and axiom generation.

arXiv ID: 2609.26029 / 要約の誤りについて