科学の課題から実行できるAI解析システムを作る枠組み
LabFactory: Building and Evaluating Executable AI Labs
この論文をやさしく読む
ひとことで言うと
科学者の依頼から、後で独立に実行・採点できる解析システムをAIに作らせる枠組み。
何に役立つ?
科学向けのAIシステムを、成果物そのものを別環境で試す形で評価する際の参考になる。
この研究の面白いところ
構築担当の説明ではなく、納品されたプログラムを未使用データで別ホストが実行して採点する。
どこまで分かった?
結果は選ばれた28件の構築例と設定された33の小課題に関するもの。任意の科学課題で同じ成功率になるとは示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
科学的な課題は求める能力を指定するが、実現には、データの取得、表現の設計、モデルの学習、道具の実装、推論時の使い方の決定など、その課題に合わせた計算システムの構築が必要なことが多い。LabFactoryは、AIの構築担当が科学的な依頼文を、実行可能なAIラボへ変える枠組みである。AIラボは、モデル、知識資源、道具、制御器を固定した入出力形式の後ろに統合した、課題専用の解決システムである。構築担当は使用量を計測する作業環境でラボを開発・梱包する。その後、別のホストが、渡された成果物を未使用の入力で実行し、正解ラベルを解決システムの入力から分離したまま、課題の手順に従って出力を採点する。これにより、構築担当の進捗説明ではなく、渡されたシステム自体を評価対象にする。 分子・ゲノム予測、生理信号、臨床判断支援、生物医学の文章など、科学的な課題7分類にわたる28件の選ばれた構築例を記録する。渡されたラボは、別ホストでの実行において、33の小課題すべてで設定された基準値を上回った。10件には構築中に学習した予測モデルが含まれ、残りは固定された基盤LLMの周囲に検索システム、実行可能な解析環境、道具を使う作業手順を組み合わせた。これらは、AIエージェントが科学的な依頼文から、構築後も呼び出し、調査し、検証できる動作するラボまで作れることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Scientific tasks specify a desired capability, but realizing it often requires building a computational system tailored to the task---acquiring data, designing representations, training models, implementing tools, and deciding how they are used at inference. We present LabFactory, a framework in which an AI builder turns a scientific brief into an executable AI lab: a task-specific solver that integrates models, knowledge resources, tools, and a controller behind a fixed interface. The builder develops and packages the lab in a metered workspace; a separate host then executes the delivered artifact on held-out inputs, with reference labels kept outside the solver's input interface, and scores its outputs under the task's protocol. This makes the delivered system, rather than the builder's account of its progress, the object of evaluation. We document 28 selected constructions across seven scientific task categories---from molecular and genomic prediction to physiological signals, clinical decision support, and biomedical text---whose delivered labs exceeded their configured reference values on all 33 subtests under host-side execution. Ten contain predictive models fitted during construction; the others assemble retrieval systems, executable analysis environments, and tool-driven workflows around a fixed platform LLM. Together they show that an AI agent can carry a scientific brief all the way to a working lab that can still be invoked, inspected, and checked after construction ends.
arXiv ID: 2609.28697 / 要約の誤りについて