arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

自律型AI研究システムScientistTwo

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Parthasarathy Ranganathan, Tomas Pfister

この論文をやさしく読む

ひとことで言うと

研究課題から仮説、実験、論文とコードまでを複数のAIエージェントで進める自律研究の枠組みです。

何に役立つ?

実験の比較基準の作成、検証、アブレーションの反復などを自動化する研究支援の設計として参考になります。

この研究の面白いところ

研究を生成するだけでなく、模擬査読と反論のループを組み込んで方法を修正します。採択済みの機械学習論文を基準に能力を評価しています。

どこまで分かった?

人の研究を上回る性能や公開可能な論文を作れるというのは著者らの報告です。査読点の比較はAI評価エージェントによるもので、生成論文が人の査読を通過した証拠とは区別されます。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

科学的発見とは、既存の知識の境界を見つけ、未踏の領域へ進む能力によって定義される。科学におけるAIの究極的な構想は、問題を起点とする自律研究である。すなわち、人間の専門家から根本的な課題を与えられたAIが、科学の状況を自ら調べ、理論的・実証的なボトルネックを見つけ、知識のフロンティアを体系的に広げる。本論文では、この構想の実現を目指す、完全自律型のマルチエージェント枠組みScientistTwoを導入する。 ScientistTwoは初期問題を入力として、最先端のベースラインを構築し、新しい仮説を定式化し、専門化したエージェントを調整して、人間の介入なしに発見サイクル全体を進める。さらに、多様なデータセットと指標を使って実験を行い、自動アブレーション研究によって方法を改良し、閉ループの模擬査読・反論エンジンで研究結果を検証する。 ScientistTwoの能力を人間の科学的成果の最高水準と比較するため、ICLR、ICML、NeurIPSなどのトップレベル会議で採択された論文を対象にベンチマークした。その結果、ScientistTwoは専門家水準で出版可能な論文と、完全に検証され実行可能なコードベースを自律的に生成したとする。解法は一貫して人間の最先端モデルを上回り、自動AI査読エージェントによる平均査読評価でも人間が書いた論文より高い評価を得た。これらの結果から、ScientistTwoは単なる支援ツールではなく、人間の発見のフロンティアを押し広げる自律的な科学の先駆者だと論じる。プロジェクトのウェブサイトも提供している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo's capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: https://scientist-two.github.io/

arXiv ID: 2609.19644 / 要約の誤りについて