arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

258の認知実験でAIと人間の判断を比べるCogGym

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum

この論文をやさしく読む

ひとことで言うと

AIが正解できるかだけでなく、人間と同じように判断するかを、認知科学の実験で比べる研究です。100論文の258実験を共通形式に整理し、50モデルを評価しています。

何に役立つ?

人間の判断をモデル化したい場合に、どの課題や入力形式でAIの応答がずれるかを調べる基盤になります。一般的な能力試験の点数とは別の視点でモデルを評価できます。

この研究の面白いところ

新しく大きいモデルは人間に近づく一方、その改善は数学やコードのベンチマークほど速くありません。人間同士の一致の安定性と比較しても、モデルの再現度には差が残ります。

どこまで分かった?

初期版は常識推論に焦点を当てた実験集です。R²は人間の応答との適合度を表し、正答率ではありません。継続的な実験追加は目標であり、現在あらゆる認知能力を網羅したという意味ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人間の知能を理解しモデル化することは、人工知能(AI)と認知科学が共有する並行した目標である。AIの能力が高まる中、モデルの応答はどのような点で人間に似ており、どこで系統的に異なるのだろうか。人間が実行し考えられる課題の幅広さと多様さは、人間とモデルを大規模かつ厳密に比較するうえで課題となる。 本研究では、同じ実験試行でモデルと人間の行動を系統的に比較するため、認知科学に基づく拡張可能な統一枠組みCogGymを提案する。CogGymは、人間が関与する半自動の処理系を用いて、多様な実験パラダイムを課題に依存しないExperiment Markup Language(EML)へ標準化し、大規模でも再現可能で忠実な比較を可能にする。初期公開では、人間の常識推論に焦点を当てた100論文から258の認知実験を選定・標準化し、50の大規模言語モデルを人間の応答と比較する。 大きく新しいAIモデルほど人間の判断をよく再現するという、明確なスケーリング傾向が見られた。しかし、このような日常的推論課題での改善は、数学やコーディングのような形式的推論ベンチマークでの向上よりかなり遅い。モデルと人間の適合度は、人間データの折半信頼性であるR²=0.93(テキスト)、0.95(画像)、0.92(動画)を大きく下回る。最良のモデルでもR²は、テキスト実験で0.59、画像で0.58、動画で0.43だった。 CogGymは、新たな認知科学実験を継続的に取り込み、モデルの行動が人間に似る点、系統的に異なる点、モデルと実験の進展に伴うその変化を明らかにする、生きた評価枠組みとなることを目指す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.

著者のコメント

Project website -- https://coggym.org

arXiv ID: 2609.21259 / 要約の誤りについて