タンパク質結合体の配列と全原子構造を整合させて設計する
CODesign: Consistency from Data to Trajectory in All-Atom Protein Binder Co-Design
この論文をやさしく読む
ひとことで言うと
標的に結合するタンパク質を設計する際、アミノ酸配列と原子の配置が互いに合うよう、データと生成手順の両方を工夫した方法です。
何に役立つ?
考えられる用途は、タンパク質やリガンドを標的とする結合体候補の計算機設計です。要旨で報告される成功率は計算機上の評価で、実験室で結合を確かめた成功率とは区別が必要です。
この研究の面白いところ
配列と構造を同時に出力するだけでは不十分と考え、約105,000の二量体データと、生成後に配列・側鎖を調整する再サンプリングを組み合わせています。
どこまで分かった?
実証はin silico、すなわち計算機上の評価です。70.9%の性能向上について、要旨には指標の詳細や相対比とポイント差の区別がありません。コードなどは公開予定とされています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
タンパク質の新規設計における中心的な課題は、互いに整合する、もっともらしい構造と配列を生成することである。すなわち、設計した各配列が意図した構造へ折り畳まれ、その構造がその配列を受け入れられる必要がある。相互依存する異なる情報形式のモデル化を切り離す典型的な2段階設計手法に比べ、共同設計モデルは配列と構造を同時に生成することで、両者の整合性を改善する。しかし、単純に同時生成するだけでは整合性は保証されない。 この課題に対処するため、CODesignの枠組みを提案する。整合性の観点で蒸留した約105,000の二量体を生成し、データの整合性を改善する。さらに、配列、主鎖構造、局所的な原子配置の同時分布を捉えるマルチモーダルな共同フローモデルと、配列および側鎖を反復的に改善する、整合性を考慮した共同再サンプリング戦略によって、整合性を高める。 実験では、CODesignはタンパク質標的とリガンド標的の両方の結合体設計で、計算機上の成功率が最も高く、最先端の性能を達成した。アブレーション実験からは、蒸留したデータセットによって性能が70.9%向上することも示された。さらに、提案した再サンプリング機構によって、無視できるほど小さい追加計算コストで性能を一段と改善できる。コード、モデルの重み、新しいデータセットは、すべてオープンソースで公開する予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The central challenge in de novo protein design is generating plausible, mutually compatible structures and sequences, such that each designed sequence folds into its intended structure and the structure accommodates that sequence. Compared to typical two-stage design methods, which decouple the modeling of the interdependent modalities, co-design models improve the cross-modal consistency by jointly generating sequences and structures. However, naively generating sequences and structures simultaneously does not ensure their consistency. To address this challenge, we propose CODesign framework. We improve data consistency by generating approximately 105,000 consistency-distilled dimers. We further promote consistency through a multimodal joint flow model that captures the joint distribution of sequences, backbone structures, and local atomic configurations, together with a consistency-aware joint resampling strategy that iteratively refines sequences and side chains. Experiments show that CODesign achieves state-of-the-art performance with the highest in silico success rates on both protein- and ligand-target binder design. Ablation studies also demonstrate our distilled dataset increases performance by 70.9%, which can be further improved by our proposed resampling mechanism with negligible additional computational cost. Code, model weights and the new dataset will be completely open-source.
arXiv ID: 2610.01773 / 要約の誤りについて