arXiv論文メモ
新着一覧
cs.DL / cs.AI / cs.SE · 査読状況未確認

論文と保存済みソースコードをWikidataで結び付ける

From Code Archival to Knowledge Graph: Bridging Software Heritage, COAR Notify and Wikidata

Camillo Carlo Pellizzari di San Girolamo and Francesco Tosoni

この論文をやさしく読む

ひとことで言うと

論文で使われたソフトウェアを見つけやすくするため、論文の識別子と保存コードの情報をWikidata上で結ぶ仕組みです。

何に役立つ?

研究ソフトウェアの発見や、論文とコードの対応確認を支える基盤になります。実際に人の確認を経て4,182件のソフトウェア項目を追加したと報告しています。

この研究の面白いところ

論文の情報とソフトウェア自体の情報を別のプロファイルで扱い、保存コードの内容に対応するSWHIDも利用します。単にURLを集めるだけでなく、知識グラフ上の関係として整備しています。

どこまで分かった?

収集対象は指定された学術誌と再現性報告です。4,397組の収集、既存82件、新規4,182件はそれぞれ異なる数です。COAR Notifyとの接続は入力形式の対応を示したもので、継続稼働する連携を実証したとは述べていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ソフトウェアは独立した重要な科学的対象であるにもかかわらず、ソースコードと学術記録との検証済みのつながりは、Linked Open Data(LOD)の網の中にほとんど存在しない。そのため、保存された成果物は意味に基づく発見から切り離されている。本論文では、論文とソースコードの関係が明示され、編集上の検証を受けている情報源から、出版物とリポジトリの組を収集、検証、モデル化する、一貫した照合処理系を提示する。情報源は、ソフトウェア中心の学術誌JOSS、SoftwareX、IPOLと、SIGMOD Availability and Reproducibility Initiative(ARI)の再現性報告である。これにより、精選された4,397組の〈DOI、リポジトリURL〉からなるコーパスを得た。 Wikidataのクラスに基づき、schema.orgとCodeMetaの語彙に整合させた、2つの異なるアプリケーションプロファイルを設計する。一方は学術論文用、もう一方はソフトウェア実体用である。この設計上の分離により、2つの粒度でルールに基づく照合が可能になる。すなわち、軽量なインラインの出版物参照、またはSoftware Heritageの内容アドレス識別子SWHIDを備えた、独立した主要対象としてのWikidataソフトウェアノードである。 Wikidataに対する読み取り専用の照会では、収集したリポジトリのうち既にモデル化されていたものは82件だけだった。その後、人が確認したバッチによって、各論文と相互にリンクする4,182件の新しいソフトウェア項目が作成された。さらに、著者らが開発しているものではない外部の取り組みである、新興のCOAR Notifyプロトコルのペイロードが、本処理系の入力形式にそのまま対応することを示す。したがって、同じバックエンドが将来、継続的な情報拡充の流れにも利用できる可能性がある。本研究の中心的な貢献は、Wikidataを学術記録と保存済みソースコードの接続点にする一対のアプリケーションプロファイルである。すべてのコード、アプリケーションプロファイル、収集データセットを公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Software is a first-class scientific object, yet validated links between source code and the scholarly record remain largely absent from the Linked Open Data (LOD) cloud, isolating archived artefacts from semantic discovery. This paper presents an end-to-end reconciliation pipeline that harvests, validates, and models publication-to-repository pairs from sources where the link between a paper and its source code is explicit and editorially verified: the software-centric journals JOSS, SoftwareX, and IPOL, together with the reproducibility reports of the SIGMOD Availability and Reproducibility Initiative (ARI). This yields a curated corpus of 4,397 $\langle$DOI, repository-URL$\rangle$ pairs. We design two distinct application profiles grounded in Wikidata classes (one for scholarly articles, one for software instances) aligned with the schema.org and CodeMeta vocabularies. This architectural separation enables rule-based reconciliation at two granularities: lightweight, inline publication references or standalone, first-class Wikidata software nodes equipped with SWHIDs, Software Heritage's content-addressed identifiers. A read-only lookup against Wikidata shows that only 82 of the harvested repositories were already modelled there; human-reviewed batches have since created 4{,}182 new software items cross-linked to their articles. We further show that payloads of the emerging COAR Notify protocol, an external effort we do not develop, map natively onto our input format, so the same backend could later serve a live enrichment stream. Our core contribution is a pair of application profiles that turn Wikidata into a connector between the scholarly record and archived source code; we openly release all code, application profiles, and harvested datasets.

著者のコメント

15 pages, 4 figures. Accepted at the 7th Wikidata Workshop (Wikidata 2026), co-located with ISWC 2026. Open-source pipeline and code available at https://github.com/ftosoni/swh-wd-reconciliation

arXiv ID: 2609.21667 / 要約の誤りについて