arXiv論文メモ
新着一覧
cond-mat.mtrl-sci · 査読状況未確認

論文の図と文章から検証付きX線回折データを構築する

ERAF4XRD: A multimodal agentic framework for constructing validated experimental X-ray diffraction databases from scientific literature

Afnan Mostafa, William Ratcliff, Simon J. L. Billinge, Niaz Abdolrahim

この論文をやさしく読む

ひとことで言うと

論文中に散らばるX線回折の図と実験条件を結び付け、出典と照合したデータベース用の記録を自動で作ります。

何に役立つ?

過去の実験結果を材料研究やAI学習で再利用できる構造化データにする用途が示されています。

この研究の面白いところ

図を見つける精度だけでなく、対応付けた最終記録を別のエージェントと人手で評価し、根拠の有無まで調べています。

どこまで分かった?

根拠のないメタデータがなかったのは、人手で評価した443フィールドの範囲です。全出力や今後の文献で誤りが出ないという保証ではなく、再現率90.7%からも抽出漏れは残ります。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

科学文献には数十年にわたる実験測定が含まれるが、現代のAIやデータ駆動型研究のための構造化データとして利用することは依然として難しい。情報の多くは図、キャプション、本文、表に分散しており、再利用の前に、実験データとその文脈を特定し、結び付け、検証する必要がある。 本研究では、科学出版物から検証済みのX線回折(XRD)記録を再構築する、完全自動の画像・文章マルチモーダルかつマルチエージェントの枠組みERAF4XRD(Experiment Reader Agentic Framework for X-Ray Diffraction)を導入する。文書のダウンロードと選別、XRD図の特定、メタデータの抽出と対応する実験データへの紐付けを行い、独立した検証エージェントで出力を出典の証拠と照合する。 3150枚の候補図を含む273件の科学出版物を人手で整理したベンチマークにおいて、XRD図の識別で最大98.7%の正解率を達成し、22フィールドにわたる1400個のメタデータ値を生成した。最終的に検証された記録の独立した人手評価では、適合率98.5%、再現率90.7%であり、評価した443フィールドには根拠のないメタデータは見られなかった。情報抽出にとどまらず、紐付けられた実験記録の再構築と検証へ進むことで、ERAF4XRDは公表済み科学情報をAIやデータ駆動科学のための機械可読な実験データセットへ変換する自動的な方法を確立する。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-17 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The scientific literature contains decades of experimental measurements that remain difficult to access as structured data for modern AI and data-driven research. Much of this information is distributed across figures, captions, text, and tables, requiring experimental data and their context to be identified, connected, and verified before they can be reused. Here we introduce ERAF4XRD (Experiment Reader Agentic Framework for X-Ray Diffraction), a fully automated multimodal (i.e., image and text), multi-agent framework that reconstructs validated X-ray diffraction (XRD) records from scientific publications. ERAF4XRD downloads and screens documents, identifies XRD figures, extracts and links metadata to the corresponding experimental data, and validates outputs against source evidence using an independent validation agent. On a manually curated benchmark of 273 scientific publications containing 3,150 candidate figures, ERAF4XRD achieved up to 98.7% accuracy for XRD figure identification and generated 1,400 metadata values across 22 fields. Independent manual assessment of the final validated records yielded 98.5% precision and 90.7% recall, with no unsupported metadata observed among 443 evaluated fields. By moving beyond information extraction to the reconstruction and validation of linked experimental records, ERAF4XRD establishes an automated approach for transforming published scientific information into machine-readable experimental datasets for AI and data-driven science.

arXiv ID: 2609.18583 / 要約の誤りについて