arXiv論文メモ
新着一覧
stat.ME / math.ST / stat.AP / stat.TH · 査読状況未確認

打ち切りのある生存データの長期予測を比較

Long-Term Tail Modeling in Survival Analysis via Extended Generalized Pareto Distributions

Eduardo Janotti, Lígia Henriques-Rodrigues, Antonio Carlos Pedroso de Lima

この論文をやさしく読む

ひとことで言うと

途中で観察が終わる生存データから、長期の生存や再発の傾向を外挿する際、模型の選び方が結果にどう影響するかを調べた研究です。

何に役立つ?

打ち切りのある医療データで長期予測を行う際、裾指数や尺度の推定に適した模型を比較する手掛かりになります。

この研究の面白いところ

三つの摂動分布を同じ当てはめ手順で比較し、標本内の適合が似ても長期予測は大きく変わると示しています。

どこまで分かった?

結果はシミュレーション条件と膀胱がん再発・心不全生存の二つの適用例に基づきます。長期外挿には模型の指定による不確実さが残ります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

右側打ち切りのある生存データを対象に、特に裾の推論と長期的な外挿を重視した、拡張一般化パレート模型の一群を提案する。この枠組みは極値理論と生存分析を統合し、一般化パレート分布と、単位区間上の柔軟な摂動分布を組み合わせる。摂動分布には、パラメトリックなベータ模型、ベルンシュタイン多項式による推定、ヒストグラムによる推定という三つの指定を考える。直接比較できるよう、三つの模型すべてを、Kaplan–Meier推定に基づく擬似観測値で右側打ち切りに対応させた、共通の反復手順で当てはめる。 モンテカルロ研究では、裾指数、打ち切りの程度、標本数、模型の複雑さを変えて、有限標本での性能を評価する。結果には柔軟性と安定性の兼ね合いが現れ、ベータ指定は一般に裾指数の推定で最も良く、ヒストグラム推定は打ち切りが少ない場合の尺度推定で特に良い。ベルンシュタイン推定は中間的な性能で、標本数と打ち切りにより敏感である。膀胱がんの再発データと心不全の生存データへの適用では、標本内での適合が非常に似た模型でも、裾指数の推定と長期外挿は大きく異なり得ることが分かった。これらの結果は、打ち切りのある生存データを拡張一般化パレート模型で外挿するとき、摂動分布の指定が重要であることを強調する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We propose a class of extended generalized Pareto models for right-censored survival data, with particular emphasis on tail inference and long-term extrapolation. Our framework integrates extreme value theory and survival analysis, combining a generalized Pareto distribution with a flexible perturbation distribution on the unit interval. We consider three perturbation specifications: a parametric Beta model, a Bernstein polynomial estimator, and a histogram-based estimator. To facilitate direct comparison, all three models are fitted using a unified iterative procedure adapted to right censoring through Kaplan-Meier-based pseudo-observations. A Monte Carlo study evaluates finite-sample performance across different tail indices, censoring levels, sample sizes, and model complexities. The results reveal a trade-off between flexibility and stability: the Beta specification generally performs best for tail-index estimation, whereas the histogram estimator performs particularly well for scale estimation under low censoring. The Bernstein estimator shows intermediate performance and greater sensitivity to sample size and censoring. Applications to bladder cancer recurrence and heart-failure survival data show that models with very similar in-sample fits can nevertheless produce markedly different tail-index estimates and long-term extrapolations. These findings emphasize the importance of perturbation specification when extended generalized Pareto models are used for survival extrapolation under censoring.

arXiv ID: 2609.28321 / 要約の誤りについて