arXiv論文メモ
新着一覧
cs.LG / cs.CE · 査読状況未確認

逆問題の生成モデルが事後分布を再現できるか評価

PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers

Jiachen Yao, Zi-Siang Hsu, Xi Deng, Aditi Gupta, Xin Ju, Sally M Benson, Gege Wen, Anima Anandkumar

この論文をやさしく読む

ひとことで言うと

逆問題の生成モデルを、一つのもっともらしい解だけでなく解の分布全体で評価するベンチマークです。4種類の物理に基づく逆問題を扱います。

何に役立つ?

同じ観測に複数の解が合う場合、推定の不確かさが適切かを評価できます。点推定が正確でも分布が崩れる問題を見つけるための枠組みです。

この研究の面白いところ

計算の重い既存手法で参照事後分布を作り、平均・分散・分布間距離・周波数構造を5指標で比較します。生成時のノイズや誘導の重みが分散の調整に重要としています。

どこまで分かった?

評価は選ばれた4課題と参照分布に基づきます。すべての科学的逆問題を代表するという保証や、現実の任意のデータでの校正保証は要旨にはありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

科学の逆問題を解くために生成モデルが使われることが増えているが、従来の評価は、もっともらしい再構成を一つ作れるかどうかに依然として重点を置く。同じ疎な観測やノイズを含む観測と整合する解が複数ある不良設定の問題では、これでは不十分である。このような状況では、点ごとの精度が高くても、モード崩壊、不確実性に対する過信、両立しない解の平均化によって、真の事後分布を捉えられないことがある。 生成型逆問題ソルバーの分布としての正確さを評価するベンチマークPosteriorBenchを導入する。Darcy流の逆解析、Poisson方程式の源の復元、炭素回収・貯留、光輸送による材質推定という四つの物理に基づく逆問題を評価する。各課題について、棄却サンプリングやマルコフ連鎖モンテカルロ法など、計算負荷は高いが確立した手順によって高忠実度の参照事後分布を作る。これにより、最良の単一標本ではなく、解の集合全体を復元できるかを直接評価できる。 これらの参照分布を、事後平均誤差、事後標準偏差誤差、最大平均不一致、スライス化ワッサースタイン距離、半径方向に平均したパワースペクトル誤差という五つの指標と組み合わせる。これらは、点ごとの正確さ、周辺的不確実性、分布の一致、全体的な周波数特性の忠実度を評価する。ベンチマークは、疎な計測、低解像度観測、非線形順モデル、異なるノイズ水準、多峰性の事前分布に及び、分布照合と不確実性の定量化を統一した処理系で行う。 実験では、現行のソルバー全般に分布照合の大きな不足が明らかになった。一方、ニューラル演算子は解像度への頑健性を改善し、ガイダンスの重みと生成ノイズが事後分散の較正の鍵となることも示された。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, overconfident uncertainty, or averaging incompatible solutions. We introduce PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport material inference. For each task, we construct high-fidelity reference posteriors using computationally heavy but established procedures such as rejection sampling and Markov chain Monte Carlo, enabling direct assessment of whether solvers recover the full set of solutions rather than the single best sample. We pair these references with a five-metric posterior evaluation suite: posterior-mean error, posterior-standard-deviation error, maximum mean discrepancy, sliced Wasserstein distance, and radially averaged power-spectrum error. These metrics assess pointwise accuracy, marginal uncertainty, distributional alignment, and global frequency fidelity. The benchmark spans sparse sensing, low-resolution observations, nonlinear forward models, varying noise levels, and multimodal priors, with a unified pipeline for distribution matching and uncertainty quantification. Our experiments reveal substantial distribution-matching gaps across current solvers, while showing that neural operators improve resolution robustness, and guidance weights and generation noise are key to posterior-variance calibration.

著者のコメント

32 pages, 10 figures, 21 tables; the code is available at https://github.com/neuraloperator/PosteriorBench

arXiv ID: 2609.20794 / 要約の誤りについて