arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

難しい負例で化合物探索を評価するベンチマーク

TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

Surbhi Kumar, Yuhe Zhou, Varun Shiralkar, Niu Huang and Baris Coskunuzer

この論文をやさしく読む

ひとことで言うと

化合物を見分けるモデルを、簡単な囮で過大評価しないための難しい負例を使う評価セット。

何に役立つ?

創薬初期の仮想スクリーニング手法を、より厳しい条件で公平に比較するのに役立つ。

この研究の面白いところ

93標的で活性化合物と似た囮を1対40で用意し、無作為な囮での高性能が厳しい条件で大きく下がることを示した。

どこまで分かった?

ベンチマークは整理したChEMBL 35データと定めた評価手順に基づく。実際の創薬での化合物発見率までは要旨に示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

リガンドに基づく仮想スクリーニング(LBVS)は初期の創薬で候補を絞る実用的な手段だが、既存のベンチマークは無作為の負例、見分けやすい囮、限られた標的、統一されない評価手順によって性能を過大評価しうる。著者らは、見分けにくい負例を使う複数標的LBVSベンチマークTopU-LBVSを導入する。整理されたChEMBL 35の生物活性データから出発し、7つのタンパク質クラスに属する93の標的を対象とする。各標的のスクリーニング用ライブラリには、性質をそろえ構造も似せた囮を活性化合物対囮1対40の固定比率で入れる。ライブラリには約400~10,000化合物があり、単純な物理化学的性質や最近傍のフィンガープリントに頼る近道を減らすよう設計されている。3つの固定手順を用意する。TopU-LBVS-fullは93標的すべてでChEMBL*からTopUへの一般化を評価し、lowは難しい負例の分布内でデータが少ない場合のTopUからTopUへの学習を評価する。miniは7標的の小型手順で、試験時の囮だけを変えた対応する無作為囮の対照を備え、低コストな開発と無作為なChEMBL*囮とTopU囮の性能差の直接測定を可能にする。フィンガープリント、分子グラフニューラルネットワーク、その混合、現代的な分子モデルを含む10の基準手法では、無作為囮で評価した性能が、難しい負例のスクリーニングでは急激に下がった。将来のLBVSと分子表現学習を再現可能に比較できるよう、データ、固定分割、評価コード、基準実装を公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but existing benchmarks can overestimate performance through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL~35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts. TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL$^\ast \rightarrow$ TopU generalization across all 93 targets. TopU-LBVS-low evaluates low-data TopU $\rightarrow$ TopU learning within the hard-negative distribution. TopU-LBVS-mini provides a compact seven-target protocol with a paired random-decoy control that changes only the test decoys, enabling low-cost development and direct measurement of the gap between random ChEMBL$^\ast$ and TopU decoys. Across ten reference baselines spanning fingerprint methods, molecular GNNs, fingerprint hybrids, and modern molecular models, performance under random-decoy evaluation degrades sharply under hard-negative screening. We release data, fixed splits, evaluation code, and baseline implementations for reproducible comparison of future LBVS and molecular representation learning methods. Code and data are available at https://github.com/topu-benchmark/topu-lbvs and https://huggingface.co/datasets/topu-benchmark/topu-lbvs.

著者のコメント

75 pages

arXiv ID: 2609.29740 / 要約の誤りについて