曖昧な結合測定からタンパク質量を推定する手法
ProteoEM: probabilistic protein abundance estimation from iterative affinity traces
この論文をやさしく読む
ひとことで言うと
一つの測定結果が複数のタンパク質候補に当てはまるとき、確率的に分けて量を推定する方法です。
何に役立つ?
単一分子の親和性測定から、プロテオフォームの存在量や試料の組成を推定する解析に役立ちます。
この研究の面白いところ
区別できない候補は無理に一つへ決めず群として示し、観測されやすさと元試料の組成も分けて扱います。
どこまで分かった?
要旨の性能結果はシミュレーションです。候補の欠落や偏った欠測で推定が偏り、実験データでの検証が今後必要とされています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
単一分子の親和性マッピングによってタンパク質やプロテオフォームを分子単位で測定できるが、プローブの結合は不完全で非特異的なため、一つの親和性の記録が複数の分子候補と一致し得る。正確な存在量の推定には、曖昧な記録を一つの候補に決め打ちせず、重みを付けて配分する必要がある。RNAシーケンスでの転写物量推定法に着想を得て、プロテオフォームを重み付きで定量する期待値最大化の枠組みProteoEMを開発し、オープンソースのPythonパッケージとして公開した。ProteoEMは、事前に較正した固定のプローブ応答率を存在量の推定とは分けて用いながら、観測された親和性の特徴の尤度全体を保持して、各分子をすべての候補と照合する。プロテオフォームの存在量を推定し、測定では区別できないものを群として報告し、観測されやすさの違いも考慮して、観測された分子の組成と元の試料の組成を区別する。シミュレーションでは、各記録を単純な二値の判定に変える方法が大きな誤差を生んだ場合でも、ProteoEMは元の分子組成を正確に復元した。中程度で一様な較正誤差には性能があまり左右されなかったが、欠測が分子の種類に依存する場合や、参照集合にないプロテオフォームがある場合には偏りが生じた。観測されやすさが既知の場合は、観測した分子数から元の試料の組成も復元できた。ProteoEMは単一分子の親和性測定を定量解析するための公開された再現可能な枠組みを提供し、これらの結果は実験で得た分子データによる検証の必要性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Single-molecule affinity mapping enables molecular-level measurement of proteins and proteoforms, but imperfect and nonspecific probe binding makes individual affinity traces compatible with multiple molecular identities. Accurate abundance estimation therefore requires apportionment of ambiguous traces by weight rather than assignment to a single candidate. We developed ProteoEM, an expectation-maximization framework for weighted proteoform quantification, inspired by transcript abundance estimation methods for RNA sequencing and released as an open-source Python package. ProteoEM evaluates each molecule against every candidate using fixed, pre-calibrated probe-response rates held separate from the abundance estimate, while retaining the full likelihood of the observed affinity features. The framework estimates proteoform abundances, reports indistinguishable proteoforms as groups when measurements cannot separate them, and accounts for differential observation yields to distinguish the composition of observed molecules from that of the source sample. In simulations, ProteoEM accurately recovered the underlying molecular composition where approaches that reduce each trace to a hard yes/no call introduced substantial errors. ProteoEM's performance was insensitive to a moderate, uniform calibration error but was biased by informative missing data and by proteoforms absent from the reference. When observation yields were known, it also recovered source-sample composition from observed molecular counts. ProteoEM provides an open-source, reproducible framework for quantitative analysis of single-molecule affinity measurements, and these results motivate validation on experimental molecule-level data.
arXiv ID: 2609.29155 / 要約の誤りについて