arXiv論文メモ
新着一覧
cs.AI / cs.CL · 査読状況未確認

未知の規則を探索して学ぶAIを測るExplorationBench

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan

この論文をやさしく読む

ひとことで言うと

事前学習の知識だけでは解けない人工の世界で、AIが実験を通じて規則を発見できるか測るベンチマーク。

何に役立つ?

AIの探索能力や、未知の規則を見つけて保留課題に使う能力の比較に役立つ。

この研究の面白いところ

実行可能な規則で答えを厳密に確認できる二つの環境を用意し、10システムを評価した。強いシステムは未知の規則を習得したが、探索の経路による差も大きかった。

どこまで分かった?

評価はAlienCodeとAlienLogicという人工環境に基づく。実際の科学研究で同じ発見能力を示すかは要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

科学的な発見は、既知の問題の先で始まる。そのためAIシステムには、仮説を立て、実験を設計し、結果を踏まえて繰り返し調べる探索が求められる。しかし、この能力の評価は難しい。新しい仮説が本当に成り立つかを検証することと、システムが探索によって発見したのか、事前学習で得た関連知識を思い出しただけなのかを区別する必要があるためだ。本研究は、検証可能なAlien Worldsに基づくExplorationBenchを導入し、科学的探索の評価を具体的で扱いやすい枠組みにする。その規則は実行可能なので答えを正確に確認でき、慣れた知識と矛盾するため、既知の知識を思い出すだけでは解けない。ベンチマークには二つのサンドボックスがあり、AlienCodeには31の発見対象と70課題、AlienLogicには24の発見対象と70課題が含まれる。各サンドボックスは、不完全な手引き、課題ごとの環境からのフィードバック、専用のツール呼び出し形式を提供する。システムはこれらを使って環境を探索し、その後、学習時に伏せた課題を解く。10のAIシステムを評価した結果、最も強いシステムは未知の規則を習得して適用できたが、探索の経路によって性能には大きな差があり、探索を続けると改善が止まったり、以前の成果が後退したりする場合もあった。ExplorationBenchは、未知の環境を探索して本当に新しい知識を獲得・適用するAIに向けた一歩となる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Scientific discovery begins where known problems end. There, AI systems must engage in exploration: framing hypotheses, designing experiments, and iterating on the results. However, evaluating this ability is difficult: (1) how to verify whether a genuinely new hypothesis holds, and (2) how to determine whether a system has discovered it through exploration or merely recalled related knowledge from pre-training data. To this end, we introduce ExplorationBench, which turns the wicked problem of evaluating scientific exploration into a concrete and tractable framework built on verifiable Alien Worlds: their rules are executable, so every answer can be checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. The benchmark contains two sandboxes, AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks). Each sandbox provides a flawed manual, task-specific environmental feedback, and a dedicated tool-call schema. Systems use these resources to explore the sandbox, then solve held-out tasks. We evaluate 10 AI systems and find that the strongest systems can acquire and apply unfamiliar rules, while performance varies substantially across trajectories and continued exploration can stall or reverse earlier gains. ExplorationBench represents a step towards AI systems that can acquire and apply genuinely new knowledge through exploration in unknown environments.

arXiv ID: 2609.30199 / 要約の誤りについて