空間トランスクリプトームの測定点選択を比較
Benchmarking Active Spot Selection for Cost-Efficient Spatial Transcriptomics
この論文をやさしく読む
ひとことで言うと
組織内で遺伝子発現を測る点を賢く選ぶ方法が、無作為抽出より良いか比較した研究。
何に役立つ?
測定費用が限られる空間トランスクリプトーム研究で、選択方法を評価する際の参考になる。
この研究の面白いところ
二つのコホートと160設定で比較し、小さい測定予算では能動的な選択が一貫して良いとは言えなかった。
どこまで分かった?
後ろ向きの模擬評価で、学習時間は固定。結果は組織の種類や評価指標によって異なる。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
空間トランスクリプトームは組織内の位置とともに遺伝子発現を測るが、密な測定格子は費用が高く、形が似た領域を繰り返し測ることもある。多くの能動学習法は分類ラベルと独立した標本向けに開発されており、高次元で連続的な発現ベクトルや、空間的に相関する候補に適するかは明らかでない。本研究は、プールから点を選ぶ能動学習と、一様な無作為抽出を後ろ向きに比較する。 発現情報がすべて得られている二つの公開コホートを使い、候補点の発現ベクトルを隠して複数回にわたる選択を模擬した。不確実性によるMC-dropoutと時間的な出力差(TOD)、多様性によるCoreSetとTypiClustに着想を得た選択を評価する。患者単位で分けた交差検証の下、各分割の訓練用測定点の5%、10%、30%、50%を使う計160設定を、全ラベルを使う参照とともに比較した。同じ予算では選択の予定、形態から発現を予測するモデル、最適化の手順をそろえた。指標は、スライド内の遺伝子ごとのPearson相関の平均、発現クラスタの一致、MoranのIの再現度である。 HER2陽性乳がんでは、四つの能動的な方法をまとめた平均相関の無作為抽出との差は、予算5%、10%、30%、50%で順に−0.0176、−0.0117、+0.0056、+0.0057だった。皮膚扁平上皮がんでは、5%で三つの方法、10%で四つすべてが無作為抽出を下回った。HER2陽性乳がんの小さい二つの予算では、CoreSetとMC-dropoutは相関が低い一方で発現クラスタの一致度は高かったが、この傾向は皮膚扁平上皮がんでは再現しなかった。報告された固定の学習時間では、小さい予算で能動的な方法は無作為抽出を一貫して上回らず、順位は評価指標によって変わる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Spatial transcriptomics (ST) measures gene expression in tissue context, but dense capture grids can be costly and may repeatedly sample morphologically similar regions. Most active learning strategies were developed for categorical labels and independent samples. We conduct a retrospective pool-based benchmark of active learning versus uniform Random sampling for ST, where expression vectors are high-dimensional and continuous and candidates are spatially correlated. Using two fully profiled public ST cohorts, we mask candidate expression vectors and simulate multi-round selection with uncertainty-based Monte Carlo dropout (MC-dropout) and temporal output discrepancy (TOD), and diversity-based CoreSet and TypiClust-inspired selection. We compare 160 completed configurations at 5%, 10%, 30%, and 50% of the fold-wide training spot pool under patient-level cross-validation, with a separate full-label reference. Within each budget, strategies share the selection schedule, morphology-to-expression predictor, and optimization protocol. We assess mean per-gene within-slide Pearson correlation coefficient (PCC), expression-cluster agreement, and Moran's I fidelity. On HER2-positive breast cancer, pooled mean PCC differences from Random across the four active strategies were -0.0176, -0.0117, +0.0056, and +0.0057 at 5%, 10%, 30%, and 50%, respectively. On cutaneous squamous cell carcinoma (cSCC), three strategies were below Random at 5%, and all four were below Random at 10%. On HER2-positive breast cancer, CoreSet and MC-dropout had lower PCC but higher expression-cluster agreement than Random at the two smallest budgets; this pattern did not reproduce on cSCC. Under the reported fixed training horizons, the evaluated active strategies do not consistently improve on Random at small budgets, and rankings depend on the evaluation measure.
arXiv ID: 2609.27208 / 要約の誤りについて