arXiv論文メモ
新着一覧
physics.chem-ph · 査読状況未確認

先端言語モデルに分子の立体構造を理解する能力が現れる

Molecular Geometry Understanding Has Unintendedly Emerged in Frontier Large Language Models

Gregorii A. Semakin (1 and 2), Timofey V. Losev (1), Ilya V. Prolomov (1), Stepan N. Ostarkov (1), Igor V. Alabugin (3), Michael G. Medvedev (1) ((1) Zelinsky Institute of Organic Chemistry RAS, (2) HSE University, (3) Florida State University)

この論文をやさしく読む

ひとことで言うと

小さな有機分子の三次元構造を与え、言語モデルが安定な配座の順序を推定できるかを調べます。

何に役立つ?

考えられる用途は分子設計の初期検討ですが、今回の実証は配座の安定性順位という限定された課題です。

この研究の面白いところ

27分子810構造で、最近の一部モデルが力場に近い精度を示し、水素結合のように言語化しやすい相互作用との関連を分析します。

どこまで分かった?

DFTエネルギーに基づく小分子の評価です。分子設計全体の能力や、能力が意図せず生じた学習上の原因を直接証明したものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)は、自然言語で表された複雑な化学問題を解く高い能力をすでに示している。しかし、化学者が行う多くの作業には、化合物の3次元構造を理解し、それについて推論することが必要である。著者らが知る限り、こうした能力はLLMの開発者が意図したものでも、現在のモデルで検証されてきたものでもない。この理解は、化学へのAI応用で究極の目標の1つとされる、自律的な分子探索や医薬品開発に不可欠である。 27種類の小さな有機分子の810個の立体配置と密度汎関数理論(DFT)のエネルギーを用いて、現在のLLMのこの能力を試験した。2025年7月より前に公開されたモデルは、配座異性体を安定性の順に並べることに苦戦した。一方、GPT-5.6 Sol、Kimi K3、Gemini 3.6 Flashを含む最近の多くの先端モデルは、汎用力場Universal Force Fieldを上回る競争力のある精度を達成し、GPT-6 AstraとClaude Opus 5は現代的な力場GFN-FFに非常に近い成績となった。 順位付けに対するモデルの説明を分析すると、科学の言語で十分に定量化され、明確に定義された分子内相互作用である水素結合を持つ分子で、最も良い性能を示すことが示唆される。一方、最新の2モデルであるGPT-6 AstraとClaude Opus 5を除くすべてのモデルは、環ひずみなど定義の緩やかな概念が支配する分子に苦戦し、そこでは力場が優れる。このことは、科学の言語そのものがLLMに制約を課している可能性を示唆する。 特に、配座異性体を順位付けする能力は、科学、コーディング、抽象推論のベンチマーク成績と強く関連しており、関連はするものの概念的には異なる課題の訓練から、意図せず生じた能力であることが示唆される。科学的知識だけから分子構造の理解へと汎化したことで、現在の先端モデルは、AIによる分子設計に向けた現実的な出発点を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) have already shown strong capabilities in solving complex chemical problems expressed in natural language. Yet, many tasks performed by chemists require understanding of 3D structures of chemical compounds and reasoning about them -- abilities, which, to the best of our knowledge, were neither intended by LLM developers nor tested in contemporary models. Such understanding is essential for autonomous molecular discovery and drug development, which is one of the Holy Grails of AI application to chemistry. We tested this capability in modern LLMs on 810 geometries and DFT energies of 27 small organic molecules. Strikingly, while models released before July 2025 struggled to rank conformers by stability, many recent frontier models, including GPT-5.6 Sol, Kimi K3, and Gemini 3.6 Flash achieved competitive accuracy, outperforming the Universal Force Field, and GPT-6 Astra and Claude Opus 5 came very close to a modern GFN-FF force field. Analysis of models' explanations for their rankings suggests that they perform best for molecules with well-defined intramolecular interactions -- hydrogen bonds -- which are well quantified in scientific language; at the same time all models except the two newest -- GPT-6 Astra and Claude Opus 5 -- struggle with molecules governed by loosely defined concepts (e.g., ring strain), where force fields excel, suggesting that the scientific language itself might impose constraints on LLMs. Notably, a model's ability to rank conformers is strongly associated with its performance on scientific, coding, and abstract-reasoning benchmarks, suggesting that it emerged unintendedly from models training on linked but conceptually different tasks. Thanks to this generalization of mere scientific knowledge into understanding molecular structures, current frontier models offer a realistic starting point for AI-driven molecular design.

著者のコメント

28 pages, 3 figures, 6 extended data figures. Supplementary Information included. Code and data available at https://github.com/TheorChemGroup/LLMConfBench

arXiv ID: 2609.20666 / 要約の誤りについて