arXiv論文メモ
新着一覧
q-bio.QM · 査読状況未確認

がん細胞株でタンパク質量を予測する遺伝子発現を探す

Retracing the Process of Translation: Proteome-wide mapping of stable transcriptomic predictors of protein abundance in cancer cell lines

Johannes Schlüter and Alexander Schönhuth

この論文をやさしく読む

ひとことで言うと

940のがん細胞株で、8,423種類のタンパク質それぞれの量を予測するのに、どの遺伝子の発現情報が役立つかを調べています。自分自身に対応する遺伝子だけでなく、まとまりとして意味のある特徴も探しています。

何に役立つ?

タンパク質測定が足りないデータの補完や、遺伝子とタンパク質の関係について新しい仮説を作る用途が考えられます。要旨が示す成果は予測特徴と関連候補の特定であり、補完の実用効果が数値で実証されたとは述べていません。

この研究の面白いところ

タンパク質を一括で扱うだけでなく、各タンパク質について特徴選択を行い、広く共通する予測因子と文脈に依存するものを分けています。免疫やHOX標的、細胞骨格など解釈可能な遺伝子群も現れています。

どこまで分かった?

対象はがん細胞株のデータで、予測に使える関連候補を見つけた研究です。因果的な制御関係の確証や患者での臨床効果ではありません。個別の予測精度や比較手法に対する改善量は要旨に示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

遺伝子発現とタンパク質量の関係を理解することは、分子生物学およびシステム生物学の中心的な課題である。遺伝子発現が転写活動を反映する一方、タンパク質は細胞の表現型を決める機能分子である。しかし、転写後および翻訳の多くの調節段階がこの関係を複雑にしており、先行研究ではRNA量とタンパク質量の相関は弱〜中程度にとどまると報告されてきた。トランスクリプトームデータからタンパク質量を予測することは依然として難しいが、とりわけプロテオームデータが限られる、あるいは利用できない場合には、生物学的な理解を得るうえで価値のある目標である。 本研究では、940のがん細胞株にわたる8,423種類のタンパク質それぞれについて、予測に役立つ遺伝子発現特徴を特定するため、リッジ回帰に基づく大規模な特徴選択手法を適用した。著者らの知る限り、この規模でタンパク質ごとの網羅的な特徴選択を実施した初めての研究である。解析により、広く予測に役立つ遺伝子特徴と、文脈に特有の遺伝子特徴の両方が明らかになった。これらには、免疫関連遺伝子、HOX転写因子の標的、細胞骨格の構成要素など、生物学的に意味のあるモジュールが含まれていた。 モデルは、個々のタンパク質と、繰り返し現れるトランスクリプトーム予測因子のパターンの水準で解釈可能性を保ちつつ、安定した遺伝子・タンパク質関連の候補を特定した。この方法は、トランスクリプトームデータからタンパク質発現を解釈可能な形でモデル化し、タンパク質量に関連するトランスクリプトーム特徴への洞察を与える。この枠組みは、仮説の生成、不完全なデータセットにおけるタンパク質データの補完、がん生物学における転写後調節の理解の深化を支援する可能性がある。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Understanding the relationship between gene expression and protein abundance is central to molecular and systems biology. While gene expression reflects transcriptional activity, proteins are the functional molecules that determine cellular phenotypes. However, numerous post-transcriptional and translational regulatory layers complicate this relationship, and prior studies have reported only weak to moderate correlations between RNA and protein levels. Predicting protein abundance from transcriptomic data remains challenging, but it is a valuable goal for biological insight, especially when proteomic data is limited or unavailable. In this study, we applied a large-scale, Ridge regression-based feature selection strategy to identify predictive gene expression features for each of 8,423 proteins across 940 cancer cell lines. To our knowledge, this is the first work to perform such comprehensive protein-wise feature selection at this scale. Our analysis revealed both globally predictive and context-specific gene features. These included biologically meaningful modules such as immune-related genes, HOX transcription factor targets, and cytoskeletal components. The models identified stable candidate gene-protein associations that remained interpretable at the level of individual proteins and recurrent transcriptomic predictor patterns. Our approach enables interpretable modeling of protein expression from transcriptomic data and provides insight into transcriptomic features associated with protein abundance. This framework may support hypothesis generation, protein imputation in incomplete datasets, and deeper understanding of post-transcriptional regulation in cancer biology.

著者のコメント

Accepted for oral presentation at ISCB-AI 2026

arXiv ID: 2609.24004 / 要約の誤りについて