過去の指導経験を再利用する対話型数学指導
REAT: A Reflective Experience-Augmented Tutoring Framework for Multi-turn Mathematical Instruction
この論文をやさしく読む
ひとことで言うと
過去の数学指導対話から得た経験を整理し、生徒の状態に合わせて再利用する言語モデルの指導手法です。
何に役立つ?
考えられる用途は、複数往復の数学学習支援です。論文では指導対話の評価で比較手法より改善したと報告しています。
この研究の面白いところ
個々の問題の解答ではなく、問題をまたいで使える指導経験として過去のやり取りを抽出し、対話中に検索します。
どこまで分かった?
要旨は実験上の指導評価の改善を述べますが、実際の生徒の長期的な学習成果は記載していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
現在の大規模言語モデルは難しい数学問題を解けるが、その能力がそのまま効果的な指導につながるわけではない。複数エージェントの構成や微調整を使う高度な指導モデルでも、指導経験を時間とともに体系的に蓄積し再利用する仕組みがないことが多く、流動的な複数往復の対話で多様な生徒に適応しにくい。本研究は、過去の対話からの経験の蒸留と、指導中の適応的な検索を組み合わせるREATを提案する。 Observer、Critic、Mentorの複数エージェントによる蒸留の処理系が、過去の対話の流れを検討し、生のやり取りを問題に依存しない構造化された指導経験へ変える。実際の指導中には、生徒の理解状態を考慮する検索モジュールが、整理した経験を取り込み、その状態に応じた段階的な支援を与える。実験では、プロンプトだけを使う方法と教師あり微調整による方法の両方を有意に上回り、とりわけ複雑で評価の低い指導場面で改善した。蒸留した経験は、異なるモデル構造と数学データセットにも頑健に一般化した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Current Large Language Models (LLMs) excel at solving complex mathematical problems, yet this proficiency does not inherently translate into effective tutoring. While advanced LLM tutors may leverage multi-agent frameworks or fine-tuning, most still lack a mechanism to systematically accumulate and reuse pedagogical experience over time, limiting their adaptability to diverse student needs during fluid, multi-turn interactions. To bridge this gap, we propose the Reflective Experience-Augmented Tutoring (REAT) framework, which couples experience distillation from historical dialogues with real-time adaptive retrieval. Driven by a multi-agent Observer-Critic-Mentor (OCM) distillation pipeline, REAT reviews past conversational trajectories and distills raw interactions into structured, problem-agnostic pedagogical experiences. During live tutoring, a state-aware retrieval module injects these curated experiences to provide adaptive scaffolding based on the student's cognitive state. Experiments demonstrate that the proposed framework significantly outperforms both prompt-only and supervised fine-tuning (SFT) baselines, particularly in improving complex, low-scoring tutoring scenarios. Crucially, the distilled experiences exhibit robust generalization across diverse model architectures and mathematical datasets.
arXiv ID: 2609.29804 / 要約の誤りについて