arXiv論文メモ
新着一覧
stat.ME · 査読状況未確認

階層データでモデルのずれを考慮するベイズ推定

A Bayesian framework for multilevel data under model mis-specification

Widemberg S. Nobre, David A. Stephens, Alexandra M. Schmidt and Erica E. M. Moodie

この論文をやさしく読む

ひとことで言うと

階層データの分析モデルが実際の生成過程とずれている場合に、母集団の値と不確実性を推定する方法を提案した。

何に役立つ?

クラスター内で観測が関連するデータの不確実性を評価する方法の検討に役立つ。実例への適用は要旨では例示として位置付けられている。

この研究の面白いところ

クラスターと個々の単位の両方に重みを与えてベイズ・ブートストラップを拡張し、シミュレーションと3種類の実データ例で比較している。

どこまで分かった?

シミュレーションでの良好な性質は部分交換可能な系列が得られる条件に基づく。要旨では、あらゆるモデルのずれや階層データに対する保証は示されていない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

階層的な過程から生じたデータに対して、解析に用いるモデルが実際の生成過程からずれているという観点から、不確実性を定量化するベイズの枠組みを提案する。特に平均構造の関数形が合っていない場合に注目し、解析モデルとデータ生成モデルの不一致がもたらす依存性の下で、対象とするパラメータをベイズ的に推定する方法を論じる。提案法は、推定関数におけるクラスター間と個々の単位間の変動を考慮しながら、母集団レベルのパラメータを推定する半パラメトリックなベイズ手続きである。拡張ディリクレモデルの階層的な重みを使い、通常のベイズ・ブートストラップを、クラスターと単位の両方の変動を扱えるよう拡張する。 シミュレーション研究では、データ生成過程と提案モデルが、関心のある未知量に関係する部分交換可能な系列をもたらすとき、提案法が頻度論的に良好な性質を持つことが示された。ラドン、2022年の国際学習到達度調査(PISA)、結核の各データへの適用を例示する。結果は、提案法が階層モデルの各種変法に対して競争力を持つことを示し、特に信用区間の幅に大きな違いが見られた。この違いは提案法の基礎となるノンパラメトリックな仮定によって説明される。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-24 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We propose a Bayesian framework for uncertainty quantification from the perspective that the working model is mis-specified in settings of a multilevel data-generating process. We focus on settings in which the mis-specification fails to match the functional form of the mean structure, and discuss Bayesian estimation of target parameters under dependence induced by a mismatch between working and data-generating models. The proposal represents a Bayesian semi-parametric procedure aimed at estimating population-level parameters while accounting for cluster- and unit-level variation in the estimating function. The proposal extends the regular Bayesian bootstrap to account for cluster- and unit-level variation using multilevel weights from an enriched Dirichlet model. Simulation studies indicate that the proposed approach has good frequentist properties when the data-generating process and the proposed model induce a partially exchangeable sequence associated with the unknown quantity of interest. Applications to radon (Gelman and Hill, 2007), Programme for International Student Assessment 2022 (OECD, 2023), and tuberculosis (Nobre et al., 2023) datasets are presented for illustrative purposes. The results demonstrate that the proposed method is competitive with variations of multilevel models, with major differences observed in the range of credible intervals, which are justified by the nonparametric assumptions underlying the proposed method.

arXiv ID: 2609.27112 / 要約の誤りについて