arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

学習と指導の両方を扱う公開教育基盤モデル

OmniEdu: Open Foundation Models for Learning and Teaching

Hao Liang, Qihan Lin, Meiyi Qiang, Linzhuang Sun, Hengyi Feng, Mingrui Chen, Sizhe Qiu, Wentao Zhang

この論文をやさしく読む

ひとことで言うと

問題解決に加えて、学習内容の理解やつまずきの診断、段階的な指導を学ばせた公開教育モデル群。

何に役立つ?

教育支援モデルの学習データ設計や、問題解決と指導の両方を評価する際の参考になる。

この研究の面白いところ

100以上の情報源を四つの教育能力で整理し、40億から270億パラメーターまで、三群の教育ベンチマークで一貫した改善を報告する。

どこまで分かった?

提示された複数のベンチマークでの評価である。実際の教室や学習者を対象にした効果検証は要旨には示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

教育用の基盤モデルには、問題を解くだけでなく、カリキュラムの構造を理解し、学習者の困難を診断し、適切な指導支援を行うことが求められる。既存の教育用言語モデルは問題解決か個別指導の一方に重点を置くことが多く、学習データの組み合わせも能力ではなくデータ源や課題で整理されている。本研究は、初等中等教育の学習と指導に向けた公開モデル群OmniEduを提示する。指示学習用データは100以上の教育資源と一般的な指示データ源を組み合わせ、教科の能力、カリキュラムとの対応、診断的推論、教育上の行動と段階的支援の四つの能力で整理する。データ処理には、決定的なクリーニング、意味の監査と書き換え、課題別の品質採点、トークン予算内で多様性を選ぶ処理、教育用指示の割り当てを組み込む。その結果、6万9999件の例と1596万の教師あり応答トークンを得て、そのうち6万951件は教育に特化した例である。40億、90億、270億パラメーターのモデルを追加学習し、一般能力に加えて、カリキュラムとの対応、初等中等教育の問題解決、教育的な個別指導を評価した。教育向けの追加学習は、いずれの規模でも教育用ベンチマークの三群すべてを一貫して改善した。OmniEdu-27BはK12-Benchで完全一致率63.12%、F1が76.69%、MathFishで85.89%、EDUMATHで86.95%、MathTutorBenchのScaffold設定で78.74%を達成した。また、比較したモデルの中でLongTutorのTeaching平均が最も高く、3.02だった。これらの結果は、問題解決、カリキュラムの理解、指導支援にわたる教育課題へ一般言語モデルを適応させるうえで、能力間のバランスを考えて選別した教師データの価値を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capability. We present OmniEdu, an open family of foundation models for K-12 learning and teaching. Its instruction-tuning corpus combines over 100 educational resources and general instruction sources, organized around four capabilities: subject competence, curriculum grounding, diagnostic reasoning, and pedagogical action and scaffolding. Our pipeline integrates deterministic cleaning, semantic auditing and rewriting, task-specific quality scoring, token-budgeted diversity selection, and pedagogical instruction assignment. It yields 69,999 examples and 15.96M supervised response tokens, including 60,951 education-specific examples. We fine-tune 4B, 9B, and 27B models and evaluate curriculum grounding, K-12 problem solving, and pedagogical tutoring, alongside general capability. Education-oriented tuning consistently improves all three educational benchmark groups across model scales. OmniEdu-27B achieves 63.12% EM and 76.69% F1 on K12-Bench, 85.89% on MathFish, 86.95% on EDUMATH, and 78.74% in MathTutorBench's Scaffold setting. It also achieves the highest Teaching average on LongTutor among the evaluated models, at 3.02. These results demonstrate the value of curated, capability-balanced supervision for adapting general language models to educational tasks spanning problem solving, curriculum understanding, and instructional support.

arXiv ID: 2609.23088 / 要約の誤りについて