arXiv論文メモ
新着一覧
q-bio.PE · 査読状況未確認

感染伝播標本の孤立例から母集団と系統導入を推定する

Demographic inference of pathogen-infected populations from partially observed transmission forests

Matthew Hall

この論文をやさしく読む

ひとことで言うと

感染者のゲノム標本で、他の標本とつながらない孤立例も使い、抽出枠の規模や独立した系統導入を推定しようとする研究です。

何に役立つ?

考えられる用途は、調査の検出力設計、一様抽出という前提の点検、感染集団の規模などの推定です。要旨ではこれらを応用案として示しています。

この研究の面白いところ

通常は除外されがちな孤立例の比率を情報として使い、病原体の時間発展モデルではなく、標本のつながり方から尤度を構成します。

どこまで分かった?

病原体動態の仮定は置きませんが、伝播構造を根付き森とし固定数を一様ランダム抽出するなどの前提はあります。臨床集団での推定精度を実測した数値は要旨にはありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ゲノム疫学には、遺伝距離の閾値でクラスターにまとめる方法や、直接伝播した可能性の高い対を特定する方法など、標本に含まれるどの個体が伝播鎖の中で近いかを調べる手法がある。標本内に近隣個体が見つからない個体、すなわちシングルトンは、通常は分析から脇へ置かれる。本研究では、それらにも情報があると論じる。標本の2個体に関連が見つかるかどうかは、抽出枠の大きさと、そこへの独立した系統導入の回数に依存する。したがって、関連する個体と関連しない個体の割合自体が、これらの量についてのデータとなる。 伝播木と抽出枠の共通部分をラベル付き根付き森として扱い、そこから固定数の頂点を一様ランダムに抽出することで、この考えを形式化する。全小行列式版の行列木定理を用い、指定された頂点集合が独立集合となる、N頂点上のラベル付き根付きk森の個数を閉形式で導出する。そこから、所定の大きさの標本に関連する個体が一つも含まれない確率と、任意の観測されたクラスター構成の尤度を、いずれも閉形式で得る。各クラスター内部の伝播構造が分かる場合と、クラスターの大きさだけが分かる場合の両方を扱う。 この方法は従来の系統動態モデルへの当てはめより、標本調査や標識再捕獲法に近い。情報は単一標本の関連構造だけから得られ、病原体の動態に関する仮定を置かない。今後の研究の検出力計算、クラスターを伝播の集中地点と解釈する際の前提となる一様抽出仮定の検定、そして母集団特性の推定そのものという3つの応用を概説する。最後に、より柔軟な実装で緩和すべき仮定を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Genomic epidemiology has several methods for identifying which sampled individuals are linked by close proximity in the transmission chain, whether by grouping them into clusters under a genetic distance threshold or by identifying probable direct transmission pairs. Individuals found to have no sampled neighbours---singletons---are usually set aside. We argue that they are informative. Whether any two sampled individuals prove to be linked depends on the size of the sampling frame and the number of independent lineage introductions into it, and thus the balance of linked and unlinked individuals is itself data about these quantities. We formalise this by treating the intersection of a transmission tree with a sampling frame as a labelled rooted forest, from which a fixed number of nodes are sampled uniformly at random. Using the all-minors matrix-tree theorem, we derive a closed-form expression for the number of labelled rooted $k$-forests on $N$ nodes in which a specified set of nodes is independent. From this we obtain, again in closed form, the probability that a sample of a given size contains no linked individuals at all, and the likelihood of an arbitrary observed configuration of clusters, both when the transmission structure within each is known, and when only the cluster sizes are. The approach is closer to a survey, or to mark-recapture, than to conventional phylodynamic model fitting: the information comes from the linkage structure of a single sample alone with no assumptions regarding pathogen dynamics. We outline three applications: power calculations for prospective studies, a test of the uniform sampling assumption that underlies the reading of clusters as transmission hotspots, and demographic inference itself. We finally set out the assumptions that a more flexible implementation would need to relax.

arXiv ID: 2609.23624 / 要約の誤りについて