arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

置換基が分子の性質をどう変えるか学ぶデータセット

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee

この論文をやさしく読む

ひとことで言うと

分子の一部を置き換えたときの性質の変化を学習・評価するデータセットを作った。

何に役立つ?

分子モデルが局所的な構造変更の効果を理解しているか評価し、学習させるために役立つ。

この研究の面白いところ

骨格、置換基、分子のすべてで学習用データと重ならない評価集合を用意し、単純な記憶との区別を図る。

どこまで分かった?

要旨はモデルの改善を報告するが、具体的なスコアや実験室での分子合成による検証は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自然言語処理の進展により、分子を扱う大規模言語モデル(LLM)はさまざまな化学課題で高い性能を示している。しかし小さく局所的な構造変更が分子の振る舞いをどう変えるかといった、細かな構造と性質の関係はまだ捉えにくい。本研究は、分子の骨格に特定の置換基を結合させたときに生じる性質の変化を表す、置換基寄与のデータセットMolSCを導入する。手作業で注釈された生物活性記録から整理し、構造上の警告となる性質、標的ごとの生物活性、物理化学的記述子を含み、学習用に置換基単位の18万1000件の例を収録する。さらに、骨格、置換基、分子の各水準でMolSCと重ならない1541件の保留評価例からなるMolSC-Benchを提案する。実験では、既存の分子LLMやGPT-5.2、Gemini-3-Flashのような強力な非公開モデルも、置換基寄与の予測では信頼性が限られていた。一方、MolSCで学習するとこの能力が大幅に改善し、さまざまな下流の分子課題でも高い性能を達成した。これらの結果は、置換基寄与の学習が、細かな分子理解の重要な要素であることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in natural language processing have led to molecular Large Language Models (LLMs) with strong performance across diverse chemistry tasks. However, they still struggle to capture fine-grained structure-property relationships, particularly how small, localized modifications alter a molecule's behavior. To address this limitation, we introduce MolSC, a dataset of substituent contributions, defined as property changes induced by attaching specific substituents to molecular scaffolds. Curated from manually annotated bioactivity records, MolSC spans structural-alert liability, target-specific bioactivity, and physicochemical descriptors, and contains 181K substituent-level examples for training. We further propose MolSC-Bench, a held-out evaluation benchmark of 1,541 examples disjoint from MolSC at the scaffold, substituent, and molecule levels. Our experiments show that existing molecular LLMs and strong proprietary models such as GPT-5.2 and Gemini-3-Flash show limited reliability in substituent contribution prediction. In contrast, training on MolSC substantially improves this ability and achieves strong performance across diverse downstream molecular tasks. These results highlight substituent contribution learning as a key component of fine-grained molecular understanding.

著者のコメント

Accepted to EMNLP 2026 Main Conference

arXiv ID: 2609.23073 / 要約の誤りについて