GRE学習でAIと人の個別指導の効果を比較
StudentBench: AI and human tutoring yield equivalent GRE learning gains
この論文をやさしく読む
ひとことで言うと
GRE対策でAI個別指導と人の指導を比較し、学習効果が統計的に同等だったと報告した。
何に役立つ?
考えられる用途は、AI指導の評価や設計である。要旨で実証したのはGREの数学・言語問題での学習効果である。
この研究の面白いところ
2383人の比較と専門家による2028組の教材評価を行い、教え方と費用の違いを検討した。
どこまで分かった?
数学指導で示された返信速度などとの関係は相関であり、因果効果とは示されていない。GRE以外への一般化も要旨からは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人工知能には人の能力を補う大きな可能性があるが、最先端の研究は主にモデルそのものの能力向上に注目してきた。本研究は、AIの教え方を評価する仕組みと公開プラットフォームStudentBenchを導入し、学生とAIの間で交わされた17万5000件超のメッセージを含む大規模なデータを集め、言語モデルによる学習効果が人の個別指導と同等かを調べる。 StudentBenchを用い、AIによる指導、人による指導、指導なしのいずれかを受けた2383人について、GREの数学と言語の問題での学習効果を測った。AIによる指導はGREの学習効果で専門家の人による指導と統計的に同等であり(p=0.015)、7領域中5領域では、最も成績の良いAI指導者が平均で人の指導者を上回った。第二の研究では、専門家の人間の指導者が、言語モデルの作った授業計画と練習問題を、評価基準に沿って2028組の比較で評価した。 二つの研究により、AI指導者間の違いを、授業計画、練習問題の作成、対話での教え方、費用、参加度の五つの観点で分けて捉えた。あるAI指導者は、人による指導と同等の学習効果(p=0.044)を、918分の1の費用で達成した。学習効果の百分率1ポイント当たりの費用は、AIが0.0052米ドル、人が4.81米ドルだった。GRE数学の指導では、AIの返信が速いほど学生の発言が多く、発言が多いほど正答した練習問題が多く、正答した練習問題が多いほど学習効果が大きいという相関があった(いずれもp<0.002)。StudentBenchは https://studentbench.org で無料公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.
著者のコメント
47 pages, including references and appendices. Project site: https://studentbench.org. GitHub: https://github.com/Handshake-AI-Research/studentbench
arXiv ID: 2609.28470 / 要約の誤りについて