機能に関わる特徴を指定して非構造タンパク質領域を生成
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
この論文をやさしく読む
ひとことで言うと
決まった立体構造を取りにくいタンパク質領域を専門に生成し、機能と関係する特徴を指定して配列を調整する研究です。
何に役立つ?
IDRの配列候補を、機能に関わる特徴の組み合わせから設計する計算手法として役立ちます。細胞内局在や転写活性の改善は予測上の結果として報告されています。
この研究の面白いところ
全長タンパク質ではなく予測IDRだけでモデルを学習し、疎なオートエンコーダーの特徴を強化学習の報酬にします。異なる機能に対応する特徴を1配列に組み合わせる考え方です。
どこまで分かった?
90%という値は8課題で対象30特徴を活性化した割合の平均で、生物学的な機能達成率ではありません。要旨には細胞実験による機能検証は記載されておらず、他のタンパク質設計への拡張は可能性として述べられています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
タンパク質の天然変性領域(IDR)は、転写制御、シグナル伝達、細胞内局在などの細胞過程で中心的な役割を担うが、機能を指定した設計は依然として難しい。構造に基づく設計法はIDRに容易には適用できず、既存のタンパク質言語モデルは全長の配列で学習されるため、折り畳まれたドメインに偏った事前分布を学ぶ。本研究では、AlphaFold Databaseから整備した5400万件の予測IDRのデータセットIDiom-DBで学習した、自己回帰型タンパク質言語モデルIDiomを提案する。IDiomは、天然IDRの組成、配列パターン、モチーフ、予測される非構造性を再現する多様な配列を生成する。 機能に関連する配列パターンを制御するため、疎なオートエンコーダーの特徴を用いた強化学習RL-SAEも導入する。これは、指定した特徴集合を活性化する配列の生成に報酬を与える事後学習法である。8つのIDR設計課題にわたり、RL-SAEの配列は、対象とした30特徴の平均90%を活性化し、活性化ステアリングの24%を上回った。RL-SAEは、ステアリングや教師ありファインチューニングと比べ、生成IDRの予測上の細胞内局在と転写活性を改善し、異なる生物学的機能に関連する特徴を1本の配列内で組み合わせられることを示す。 このようにIDiomとRL-SAEは、機能に関連する配列特徴を明示的に制御することで、解釈可能で組み合わせ可能なIDR設計を実現する。より広くは、解釈可能な特徴が有用な設計目標になる他のタンパク質設計にもRL-SAEを拡張できる可能性がある。コードはhttps://github.com/rotskoff-group/idiom で公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Intrinsically disordered protein regions (IDRs) play central roles in cellular processes such as transcriptional regulation, signal transduction, and subcellular localization, yet their functional design remains challenging. Structure-based design methods do not readily apply to IDRs, and existing protein language models are trained on full-length protein sequences, thus learning a prior that is biased towards folded domains. Here, we present IDiom, an autoregressive protein language model trained on IDiom-DB, a dataset of 54 million predicted IDRs curated from the AlphaFold Database. IDiom generates diverse sequences that recapitulate the composition, patterning, motifs, and predicted disorder of natural IDRs. To control function-associated sequence patterns, we also introduce reinforcement learning with sparse autoencoder features (RL-SAE), a post-training method that rewards the generation of sequences that activate specified feature sets. Across eight IDR design tasks, RL-SAE sequences activate, on average, 90% of 30 targeted features, compared to 24% for activation steering. We demonstrate that RL-SAE improves the predicted subcellular localization and transcriptional activity of generated IDRs compared to steering and supervised fine-tuning, and enables features associated with distinct biological functions to be combined within individual sequences. Thus, IDiom and RL-SAE enable interpretable and composable IDR design through explicit control of function-associated sequence features. More broadly, RL-SAE could extend to other protein design settings where interpretable features provide useful design targets. Code is available at https://github.com/rotskoff-group/idiom.
arXiv ID: 2610.02189 / 要約の誤りについて