arXiv論文メモ
新着一覧
cs.LG / cs.CV · 査読状況未確認

少量のデータで天然変性タンパク質を予測する学習法

MT-ProtBERT: Multi-task Learning ProtBERT for Intrinsically Disordered Proteins Classification with Scarce Data

Jian Sun, Kingshuk Ghosh, Lilianna Houston, Mohammad H. Mahoor

この論文をやさしく読む

ひとことで言うと

データが少ない天然変性タンパク質の性質を、配列学習と生化学的な補助課題を組み合わせて予測した。

何に役立つ?

リン酸化部位やタンパク質のコンパクトさを、少量の配列データから推定する研究に役立つ。

この研究の面白いところ

マスク言語モデルと生化学的課題を同時に学習し、二種類の予測課題のすべてで比較対象のPARROTを上回った。

どこまで分かった?

評価は要旨に挙げられた小規模データセットと二つの課題であり、実験で立体構造を直接決定した研究ではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

天然変性タンパク質(IDP)は、折りたたまれたタンパク質と異なり、動的で安定した三次元構造を持たず、似たタンパク質同士でも配列の類似性が低い。その多様な機能に有利な構造の不均一性は、従来の実験手法で立体構造を決めることを難しくする。実験の難しさと配列類似性の低さからデータが乏しくなり、生物学や進化の理解に関係するIDPの類似性や相違の分類・検出が難しい。 本研究は、データが少ない状況に合わせたProtBERTのマルチタスク拡張、Multi-task ProtBERT(MT-ProtBERT)でこの課題に取り組む。動的な窓のマスキング、多尺度の一次元畳み込み分類器(MS-Conv1D)、マスク言語モデルと生化学的知識に基づく課題を共同で最適化する補助目的を組み合わせた。限られたデータの二つの課題で評価した。一つは短い配列と小規模データセットを使うセリン・トレオニン・チロシン(S/T/Y)のリン酸化部位予測、もう一つは684配列と530配列の二つの小規模データセットによるタンパク質のコンパクトさの予測で、典型的な変性領域に近い長さの配列も含む。MT-ProtBERTは、すべての課題でRNNベースのIDP専用モデルPARROTを一貫して上回った。結果は、自己教師あり学習、生化学の知識に基づく課題、多尺度学習の組み合わせが、データ不足下で構造を持たないタンパク質を頑健にモデル化できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Intrinsically disordered proteins (IDPs) differ from folded proteins in that they are dynamic, lack a stable three-dimensional conformation, and have low sequence similarity between similar proteins. The conformational heterogeneity of IDPs - while beneficial for their diverse functions - limits the use of traditional experimental tools to determine their conformation. The experimental difficulty, along with low sequence similarity, results in data scarcity, and makes it difficult to classify/detect IDPs that are similar or dissimilar, a task relevant to understand biology and evolution. We address this challenge using Multi-task ProtBERT (MT-ProtBERT), a multi-task extension of ProtBERT tailored for low-data regimes. MT-ProtBERT integrates Dynamic Window Masking, a Multi-Scale 1D Convolutional classifier (MS-Conv1D), and auxiliary objectives that jointly optimize masked language modeling and biochemistry-informed tasks. We evaluate this framework on two tasks under limited data: (i) phosphorylation site prediction (S/T/Y) in short sequences and small datasets, and (ii) protein compaction prediction on two small datasets (684 and 530 sequences), including sequences comparable in length to typical disordered regions. MT-ProtBERT consistently outperforms PARROT, an RNN-based IDP-specific model, across all tasks. These results demonstrate that combining self-supervised and biochemistry-informed tasks, and multi-scale learning enables robust modeling of unstructured proteins under data scarcity.

著者のコメント

14 pages, 6 figures, 12 tables

arXiv ID: 2609.25334 / 要約の誤りについて