論文の表や補足資料から知識グラフを作るArticleMiner
ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications
この論文をやさしく読む
ひとことで言うと
論文の表や補足資料にある数値の意味を、分野ごとの限定した規則と複数の解析結果から解釈し、知識グラフにする方法を提案した。
何に役立つ?
科学論文から定量的な事実をRDFなどの構造化データへ変換する仕組みを設計する際の参考になる。
この研究の面白いところ
共通の処理手順に、作業ごとの標準名や妥当性制約だけを与える構成とし、四分野163報で評価した。四作業すべてで比較手法より点推定値が良かった。
どこまで分かった?
小規模な二つのベンチマークでは結果に不確かさがある。地球化学の改善も、分野別モジュールだけの効果とは切り分けられない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
科学論文の定量的な内容は表や補足ファイルに多く含まれる。数値の意味は、見出し、説明文、単位、分析法、分野の慣習によって決まるため、表の行と列を復元するだけでは、報告された科学的事実を復元したことにはならない。意味を解釈する既存の表処理法の多くは、整った表がすでにあると仮定してから、セルや列をオントロジーの用語に対応づける。一方、論文全体から情報を取り出すシステムの多くは単一分野向けである。本研究は、その中間として、論文と補足ファイルを読み、複数の解析器と言語モデルの根拠を集めて照合する共通の処理手順を調べる。作業ごとに、人が作成した範囲の限定されたモジュールが、その分野での意味を与える。モジュールには、グラフで使える標準名、それに対応する表記、少数の導出規則と妥当性の制約、同一性を判断するキー、RDFに書き込むための対応づけを記す。これにより作業が出力できる内容を定めるが、分野の慣習をすべて列挙しようとはしない。 ArticleMinerの枠組みで、創薬化学、材料科学、機械学習、鉱物地球化学の四つのモジュールを作り、専門家が正解データを整えた新しい地球化学のベンチマークを含む、163報の論文で評価した。同じLLMを用いる少数例の提示による比較手法との比較では、四つの作業すべてでArticleMinerの点推定値が優位だった。ただし、規模の小さい二つのベンチマークでは不確かさが残る。地球化学の比較では、比較手法にも補足ファイルへのアクセスを与えたため、改善を分野固有の指示だけに帰することはできない。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 掲載先の記載あり
著者による掲載先の記載:6th International Workshop on Scientific Knowledge: Representation, Discovery, and Assessment (Sci-K), Oct 2026, Bari, Italy。出版社での独立確認は未実施です。
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Scientific papers keep much of their quantitative content in tables and supplementary files, where a number means something only through its header, caption, unit, analytical method, and the conventions of its field. Recovering the rows and columns of a table is therefore not the same as recovering the scientific fact it reports. Most semantic table-interpretation methods assume that a clean table is already available and subsequently map its cells or columns to ontology terms, whereas most publication-level extraction systems are designed for a single domain. We study a middle path: a shared process that reads a paper and its supplementary files, gathers evidence from several parsers and a language model, and reconciles that evidence, while a bounded human-authored task module for each task supplies the domain meaning. The module lists the canonical names the graph may use, the surface forms that map to them, a small set of derivation rules and validity constraints, an identity key, and the bindings used to write RDF. It defines what a task is allowed to emit; it does not try to list every convention of a field. We build four such modules (for drug-discovery chemistry, materials science, machine learning, and mineral geochemistry) in the ArticleMiner framework, and evaluate them on 163 papers, including a new geochemistry benchmark with expert-curated ground truth. In comparisons against a same-LLM few-shot baseline, the point estimates favor ArticleMiner on all four tasks, with uncertainty on the two smaller benchmarks. The geochemistry comparison also includes access to supplementary files, so its improvement cannot be attributed to domain guidance alone.
著者のコメント
Sci-K co-located with ISWC 2026, Bari (Italy)
arXiv ID: 2609.25607 / 要約の誤りについて