arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

学習者の理解と誤解を記憶し次の問題を作るAI家庭教師

CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop

Kailai He, Zhihao Wu, Linhai Zhang, Runcong Zhao, Yulan He, Jiazheng Li

この論文をやさしく読む

ひとことで言うと

正解・不正解だけでなく、何を理解し何を誤解しているかを記憶しながら、次の練習問題を変えるAI教師です。

何に役立つ?

考えられる用途は、苦手な話題や繰り返す誤解に合わせた練習支援です。進捗表示により、なぜ問題が個別化されたかを確かめる仕組みも含みます。

この研究の面白いところ

自由記述などから言語モデルが連続的な証拠を取り出し、それをベイズ的な習熟度の更新に使っています。

どこまで分かった?

68〜69%は生成問題が好まれた割合であり、学力の向上率ではありません。習熟度推定の収束は人物像シミュレーションでの結果で、要旨に長期的な学習成果の検証はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

よい指導は個人に適応する。学習者が何を知っているかを追い、なぜ間違えるかに気付き、最も役立つ次の問いを出す。これに対し、運用されている指導ツールの多くは固定した問題集を提示し、誤答を一ビットの信号として扱う。本研究では、反復的な指導ループを支える対話型エージェント教師CoLearnを提示する。学習者が練習すると、システムは習熟度と誤概念について、証拠に基づく記憶を構築する。この記憶は証拠の蓄積に応じて更新され、次の個別化された問題の生成に使われる。 CoLearnは三つの要素からなる。第一に、大規模言語モデルを連続値の観測関数として使い、ベイズ知識追跡のソフトな証拠を用いる変形によって、話題ごとの習熟度を更新する持続的な学習者状態の記憶。第二に、学習者が最も苦手とする話題と繰り返す誤概念を狙う適応的な問題生成。第三に、進捗のリアルタイム可視化とブラインドA/B比較により、個別化を見える形で検証可能にする証拠表示である。 ブラインドA/B評価では、この記憶を条件として生成した問題が、個別化されていない問題より68〜69%の割合で好まれた。また、真の習熟度を隠した人物像シミュレーションでは、エージェントの信念が学習者の真の習熟度に向かって収束する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Good tutoring adapts to the individual: it tracks what a learner knows, notices why they go wrong, and asks the next question that will help most. Most deployed tutoring tools instead serve fixed item banks and treat a wrong answer as a single bit of signal. We present CoLearn, an interactive, agentic tutor that supports an iterative tutoring loop: the learner practises, and the system builds an evidence-grounded memory of the learner's mastery and misconceptions. This memory is updated as evidence accumulates and is used to generate the next personalised question. CoLearn has three components: (i) a persistent learner-state memory that updates per-topic mastery with a soft-evidence variant of Bayesian Knowledge Tracing, where a large language model acts as a continuous observation function; (ii) adaptive question generation that targets the learner's weakest topic and recurring misconceptions; and (iii) an evidence view that makes personalisation visible and testable through live progress visualisation and blind A/B comparison. In blind A/B evaluation, questions conditioned on this memory are preferred over non-personalised ones 68-69% of the time, and in persona simulations with hidden ground-truth mastery the agent's belief converges toward the learner's true mastery.

著者のコメント

Accepted to EMNLP 2026

arXiv ID: 2609.21154 / 要約の誤りについて