arXiv論文メモ
新着一覧
cs.LG / q-bio.QM / stat.ME · 査読状況未確認

少数回のCTから病変の大きさを予測する精度と不確かさを検証

Reliability-Centered Evaluation of Sparse Longitudinal CT Lesion-Size Forecasting with Conformal Interval Calibration and Gompertz-Inspired Regularization

Lingfei Kong

この論文をやさしく読む

ひとことで言うと

過去のCT画像を多く与えれば病変の大きさをより正確に予測できるのかを、予測する区間をそろえて調べています。このデータでは、履歴を増やすだけでは改善しませんでした。

何に役立つ?

画像予測モデルを比較する際に、予測値の誤差だけでなく、予測区間が実際の値をどの程度含むかも確認する評価方法として役立ちます。治療効果や診療成績の改善を検証した研究ではありません。

この研究の面白いところ

予測期間の違いを抑えて履歴数の効果を調べ、さらに成長曲線に基づく正則化も検討しています。その項の最適な重みがゼロだったという、追加構造が役立たなかった結果も報告しています。

どこまで分かった?

対象は129人・205病変軌跡の5回観察ベンチマークです。観察回の番号で区間を固定した設計であり、同じ暦日間隔と読み替えることはできません。区間較正の被覆率改善には幅が広がる代償があり、部位による難易度差も残ります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

CTの経時的な追跡が疎で、過去の観察が数回しかないと、病変サイズの予測には制約が生じる。DeepLesionとDeep Lesion Tracker(DLT)から、同一病変の5回の観察軌跡を持つDLT由来のベンチマークを構築し、129人の患者から205本の軌跡を得た。従来型の、少数の観察から最終時点を予測する探索的解析と、主要解析として観察回の番号に基づく予測区間を固定する設計を比較した。後者はT3からT4への共通の対数変化を予測し、より早期の観察を段階的に加えるものである。予測精度、不確かさの信頼性、事後的なコンフォーマル区間較正、サブグループ別性能、ゴンペルツ曲線に着想を得た軌跡正則化を評価した。 評価した手法では点予測の精度が部分的に重なる一方、不確かさの挙動は異なった。学習用乱数シード10通りでの、評価用に取り置いたデータに対する平均RMSEは、m=1、2、3、4でそれぞれ0.4726、0.4305、0.4499、0.4513だった。平均RMSEはm=2で最小となり、追加の履歴はRMSEを改善しなかった。m=4では、未較正のCohort-Level Feature GPの被覆率が名目値95%に近かった一方、MC Dropout、Deep Ensemble、残差スケールに基づく区間は保守的だった。患者単位のコンフォーマル較正は一般に、区間が広がる代わりに、名目値に近いか保守的な被覆率をもたらした。 患者ごとにグループ化した開発段階の交差検証では、ゴンペルツに着想を得た項に対してλ*=0が選ばれた。全体集団の基準はしばしば個々の病変の変化方向に反し、予測の難しさは解剖学的なサブグループによって異なった。全体として、予測区間を統制すると、過去の観察を増やしても予測への利益は限られた。予測精度、不確かさの信頼性、軌跡の整合性は必ずしも同時に改善せず、疎な経時画像の解析ではこれらを一緒に評価すべきである。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Sparse longitudinal CT follow-up limits lesion-size forecasting when only a few prior observations are available. We constructed a five-visit DLT-derived same-lesion trajectory benchmark from DeepLesion and Deep Lesion Tracker (DLT), yielding 205 trajectories from 129 patients. We compared an exploratory conventional sparse-to-final analysis with a primary fixed visit-index horizon design predicting the common log change from T3 to T4 while progressively adding earlier observations, evaluating predictive accuracy, uncertainty reliability, post-hoc conformal interval calibration, subgroup performance, and Gompertz-inspired trajectory regularization. The evaluated methods showed partially overlapping point-prediction accuracy but distinct uncertainty behavior. Mean held-out RMSE across ten training seeds was 0.4726, 0.4305, 0.4499, and 0.4513 for m = 1, 2, 3, 4, indicating the lowest mean RMSE at m = 2; additional history did not improve RMSE. At m = 4, raw Cohort-Level Feature GP coverage was near the 95% nominal level, whereas MC Dropout, Deep Ensemble, and residual-scale intervals were conservative. Patient-level conformal calibration generally produced near-nominal or conservative coverage at the cost of wider intervals. Patient-grouped development cross-validation selected lambda* = 0 for the Gompertz-inspired term. A global population reference frequently opposed lesion-level change directions, and prediction difficulty varied across anatomical subgroups. Overall, additional historical observations provided limited predictive benefit once the prediction horizon was controlled, while predictive accuracy, uncertainty reliability, and trajectory consistency did not necessarily improve together, and should be evaluated jointly in sparse longitudinal imaging.

著者のコメント

25 pages, 11 figures

arXiv ID: 2609.21197 / 要約の誤りについて