arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

分子の環構造を文法規則の列で表して学習・生成

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi, Simone Foti, Jianmin Wang, Jure Leskovec, Tolga Birdal

この論文をやさしく読む

ひとことで言うと

分子の環や繰り返し構造を、文法の規則列として表す方法です。複雑な構造を保ちながら、通常の系列モデルで学習できる形に変えます。

何に役立つ?

分子生成や分子特性の予測に使う表現の設計・比較に役立ちます。環の種類を広く含むデータセットと網羅性の指標も提示されています。

この研究の面白いところ

高次構造を直接大きなデータ構造で処理する代わりに、生成規則へ変換します。生成時の形式的な妥当性と、環構造の表現力を両立させる狙いです。

どこまで分かった?

100%の妥当性は表現・文法の構成上の性質で、生成分子すべての合成可能性や有用な生物活性を示すものではありません。AUCの改善幅はプロービングと全体学習で異なります。要旨には実際の化学合成による評価はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

分子学習モデルは、基盤となる表現に強く左右される。しかし、標準的な系列やグラフの形式では、環構造や繰り返し現れるモチーフなどの高次の位相構造を明示的に符号化することが難しい。既存の高次表現はこうした構造を直接捉えられるが、計算負担が大きく、妥当な分子へ復号しにくいことが多い。私たちは、分子を組合せ複体へ持ち上げ、各複体を文脈自由な高次文法のもとでコンパクトな生成規則の列へ解析する、原理に基づき位相を考慮した枠組みHigher-order Grammar Representation(HGR)を提案する。高次の位相構造を規則列へ直列化することで、HGRは標準的な系列モデルに直接適合させ、位相的な表現力を保ちながら、明示的な高次符号化の計算負担を回避する。 単純な環構造に偏ったベンチマーク評価を減らすため、整備された部分集合RingDiv300kを含む118万分子の、環構造を充実させたベンチマークRingDivを構築する。また、環構造の網羅性を定量化する環多様性指標RDIを導入する。分子生成では、HGRに基づくモデルは、構成上100%の妥当性と最先端の分布整合性を独自に両立し、5つすべての生成ベンチマークでFCDが1位となった。表現学習では、HGR-FMが2つの転移プロトコルのいずれでも、7つのMoleculeNetベンチマーク全体で最高の平均AUCを達成した。最も強力なベースラインに対する改善は、プロービングで8.3 AUCポイント、全体のファインチューニングで3.3 AUCポイントだった。これらの結果は、HGRが分子生成と転移可能な表現学習に有効な、効率的な高次表現であることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Molecular learning models are strongly shaped by their underlying representations. Yet standard sequential and graph formalisms struggle to explicitly encode higher-order topology, such as ring systems and recurring motifs. Existing higher-order representations can capture these structures directly, but they are often computationally demanding and difficult to decode into valid molecules. Here, we introduce Higher-order Grammar Representation (HGR), a principled, topology-aware framework that lifts molecules to combinatorial complexes and parses each complex into a compact sequence of production rules under a context-free higher-order grammar. By serialising higher-order topology into rule sequences, HGR makes these structures directly compatible with standard sequence models, avoiding the computational overhead of explicit higher-order encodings while preserving topological expressiveness. To reduce benchmark bias towards simple ring systems, we construct RingDiv, a ring-enriched benchmark containing 1.18 million molecules, including the curated RingDiv300k subset, and introduce the ring diversity index (RDI) to quantify ring-system coverage. In molecular generation, HGR-based models uniquely combine 100% validity by construction with leading distributional alignment, ranking first in FCD on all five generation benchmarks. In representation learning, HGR-FM achieves the highest mean AUC across seven MoleculeNet benchmarks under both transfer protocols, improving on the strongest baseline by 8.3 and 3.3 AUC points under probing and full fine-tuning, respectively. Collectively, these results establish HGR as an efficient higher-order representation for molecular generation and transferable representation learning.

arXiv ID: 2610.02186 / 要約の誤りについて