arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

文章からベイズ網の構造と確率を取り出す評価データ

PRISM-BN: A Controlled Corpus and Benchmark for Text-to-Parameterized Bayesian Network Extraction

Amartya Bhattacharya, Nikhil Singh, Neeti Pokhriyal, Soroush Vosoughi

この論文をやさしく読む

ひとことで言うと

文章に書かれた関係を図として取り出す力と、その関係に正しい確率を付ける力を別々に測るデータ集です。

何に役立つ?

文章からベイズネットワークを作るAIの比較や学習に使えます。関係図がそれらしくできても、確率まで一致しているとは限らないことを評価できます。

この研究の面白いところ

ノード、状態、辺、完全な条件付き確率表を一組にして評価します。辺の復元は良好でも、親が複数ある場合も含めて確率分布を厳密に合わせるのは難しいという差が見えています。

どこまで分かった?

確率は評価用に内部で作った値で、現実の因果関係を外部データで検証した推定ではありません。高いEdge F1は条件付き評価の値です。人による確認も予備調査として位置付けられています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

確率的グラフィカルモデル(PGM)、特にベイズネットワーク(BN)は、有向の構造と確率パラメータを明示するため、ニューロシンボリックAIにとって自然な記号的出力対象となる。しかし、文章からパラメータ付きBNを作るシステムの学習には、文章とBNを対応付けた資源が必要であり、大規模なものは利用できなかった。本研究では、5分野にわたり、BNに基づく5,054件の記述を、変数、状態、有向辺、根ノードの事前確率、複数の親を持つ場合の完全な条件付き確率分布(CPD)を含む離散型の参照BNと対応付けた、統制されたコーパスPRISM-BNを導入する。事例はWikipediaを出発点とする50の骨格から派生させており、その確率は内部で構成したベンチマークの目標値であって、外部で検証された因果的な推定値ではない。 PRISM-BNは、周辺分布を先に扱う処理系PRISMを用いて構築する。この処理系は、周辺分布と局所的な同時分布を引き出し、正規化されたCPDを解析的に復元し、局所的にパラメータを再設定した部分グラフを構成する。意味に基づくノードと状態の対応付け、条件付きの構造採点、厳格な完全CPD評価を備えるベンチマークを定義する。 6つのLLM抽出器において、Node F1は0.56〜0.83、条件付きEdge F1は0.90〜0.97、CPD-KLは1.11〜3.14だった。条件付きの状態・辺の復元は一貫して良好である一方、厳格な完全CPDの一致は依然として難しい。この傾向は、GPT-5.5で独立に生成した参照でも持続し、人による予備調査でも、構造が復元可能であることと、確率の解釈がおおむね類似することが裏付けられた。PRISM-BNは、構造の復元と確率パラメータの推定を分けて評価することを可能にする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Probabilistic Graphical Models (PGMs), especially Bayesian Networks (BNs), expose directed structure and probabilistic parameters, making them natural symbolic targets for neurosymbolic AI. Yet training text-to-parameterized-BN systems requires paired text-to-BN resources unavailable at scale. We introduce PRISM-BN, a controlled corpus of 5054 BN-grounded descriptions paired with discrete reference BNs containing variables, states, directed edges, root priors, and full multi-parent CPDs across five domains. The instances are derived from 50 Wikipedia-seeded backbones, and their probabilities are internally constructed benchmark targets rather than externally validated causal estimates. PRISM-BN is built with PRISM, a marginal-first pipeline that elicits marginal and local joint distributions, analytically recovers normalized CPDs, and constructs locally reparameterized subgraphs. We define a benchmark with semantic node and state alignment, conditional structural scoring, and strict full-CPD evaluation. Across six LLM extractors, Node F1 ranges from 0.56 to 0.83, conditional Edge F1 from 0.90 to 0.97, and CPD-KL from 1.11 to 3.14. Conditional state and edge recovery remain consistently strong, whereas strict full-CPD agreement remains challenging. These trends persist with independently generated GPT-5.5 references, and a human pilot corroborates structural recoverability and similar probabilistic interpretations. PRISM-BN supports separate evaluation of structural recovery and probabilistic parameter estimation.

arXiv ID: 2609.21673 / 要約の誤りについて