arXiv論文メモ
新着一覧
physics.chem-ph · 査読状況未確認

化学・材料の実験データで最適化手法を比べるRECOB

RECOB: Reliable Benchmarking of Experimental Optimization in Chemistry and Materials Science

Zikai Xie, Jiaming Wan, Linjiang Chen

この論文をやさしく読む

ひとことで言うと

化学や材料の実験条件を効率よく探すアルゴリズムを、実際の実験データに基づいて比較する評価セットです。

何に役立つ?

どの最適化手法を実験計画に使うかを、性能と計算負担の両方から検討する材料になります。

この研究の面白いところ

実験データを学習した代理的な応答モデルだけでなく、元の実測値のリプレイでも順位を確かめています。

どこまで分かった?

比較対象は単目的14課題と多目的2課題です。学習したオラクルへの問い合わせは新しい物理実験そのものではなく、順位の再現性もこの評価範囲での結果です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

実験科学向けの最適化手法は、再現可能ではあっても実験の重要な特性を欠く人工的な関数で評価されることが多い。本研究では、化学・材料科学の物理的な実験で生成されたデータだけから構築したブラックボックス最適化ベンチマークRECOB(REliable Chem Optimization Benchmark、GitHubリポジトリ:https://github.com/XieZikai/RECOB)を提案する。化学反応、材料配合、電気化学系、連続フロープロセス、自動実験室にわたる単目的14課題と多目的2課題を含む。各課題には、決定変数、実行可能領域、物理的制約、目的関数を最小化するか最大化するか、実験データの由来を機械可読形式で示す。 連続的に問い合わせ可能な学習済みオラクルは、反復ホールドアウト検証と事前に定めた採用基準で選別する。一方、実測値テーブルのリプレイでは元の実験応答を使って評価できる。共通の対応付き評価手順で、単目的最適化10手法と多目的最適化8手法を比較した。単目的の総合順位ではHEBOが最良となり、多目的の比較ではqNEHVIが首位となった。モデルに基づく手法は概して非適応的なベースラインを上回るが、計算上の付加負担には大きな違いがある。 さらに、オラクルの独立した再学習と実測値テーブルのリプレイで、ベンチマークの信頼性を評価した。最適化手法の総合順位は、再学習したオラクル間で高い一貫性を維持し、物理的に測定された応答だけを使うリプレイでも大まかな性能の序列が保たれた。以上は、RECOBが実験に根差し、信頼性も検証されたベンチマークとして、化学・材料科学のブラックボックス最適化手法の性能差を再現可能に識別できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Optimization methods for experimental science are often evaluated on synthetic functions that are reproducible but omit important characteristics of real experiments. We introduce RECOB (REliable Chem Optimization Benchmark, Github repository: \hyperlink{https://github.com/XieZikai/RECOB}{https://github.com/XieZikai/RECOB}), a black-box optimization benchmark constructed exclusively from data generated through physical experiments in chemistry and materials science. The suite contains 14 single-objective and two multi-objective tasks spanning chemical reactions, material formulations, electrochemical systems, continuous-flow processes, and automated laboratories. Each task provides a machine-readable specification of its decision variables, feasible domain, physical constraints, objective direction, and experimental provenance. Continuously queryable learned oracles are screened using repeated holdout validation and prespecified admission criteria, while measured-table replay enables evaluation using the original experimental responses. Under a common paired evaluation protocol, we compare ten single-objective and eight multi-objective optimization methods. HEBO achieves the best aggregate single-objective rank, while qNEHVI leads the multi-objective comparison. Model-based methods generally outperform non-adaptive baselines, although their computational overhead varies substantially. We further assess benchmark reliability using independent oracle retraining and measured-table replay. Aggregate optimizer rankings remain highly consistent across retrained oracles, while replay preserves the broad performance hierarchy using only physically measured responses. Together, these results show that RECOB can reproducibly distinguish optimizer performance as an experimentally grounded and reliability-tested benchmark for black-box optimization in chemistry and materials science.

arXiv ID: 2609.20891 / 要約の誤りについて