データ科学授業での学生のLLM利用と自動生成問題の難度
Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses
この論文をやさしく読む
ひとことで言うと
データ科学の授業で学生のLLM利用を調べ、LLMが作った問題の難度表示が実際の難しさを表すか検証した。
何に役立つ?
授業で自動生成した選択式問題を使う際、LLMの難度表示をそのまま採用せず、学生の回答データで確認する判断に役立つ。
この研究の面白いところ
難度表示とLLM自身が付けたブルーム分類は強く相関したが、どちらも学生の実際の正答傾向から測った難度をほとんど予測しなかった。
どこまで分かった?
10週間、ドレクセル大学の3授業での観察結果。要旨は他大学や他分野で同じ結果になるかは示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ドレクセル大学のデータ科学の授業で、10週間の学期にわたり複数の情報源を用いた教室研究を行った。まず3つの授業で4回の調査を実施し、学生が大規模言語モデル(LLM)をどう使うかを調べた。次に、授業中の復習練習向けにLLMが生成した選択式問題の構成概念妥当性を評価した。作成した378問のうち311問を実際に使用し、得られた学生の回答7,888件から、LLMが付けた難度評価が実測の問題の難しさと一致するかを分析した。 学生のLLM利用の程度は授業間で異なり、学期が進むにつれて増えた。学生の満足度は高く、かなりの時間を節約できたと報告した一方、深い学習に役立つという認識は低下し、過度な依存傾向を挙げる学生も多かった。LLM生成問題の難度評価では、易・中・難の区分は、同じLLMが付けたブルームの教育目標分類の水準と強く相関した(スピアマンのρ=0.90)。これは両者を同時生成したことによる見かけ上の関係を反映する。しかし、どちらも実測の問題難度は予測しなかった(難度区分のρ=0.06、ブルーム水準のρ=0.02)。評価は問題の本質的な難しさよりも、構造上の書式を反映していた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
This paper presents a multi-source classroom study conducted during a 10-week quarter in data science courses at Drexel University. We first investigate the behaviors of student engagement with large language models (LLMs) using four surveys across three data science courses. Second, we evaluate the construct validity of multiple-choice questions (MCQs) generated by an LLM for in-lecture retrieval practice. Based on 378 authored questions (311 deployed, producing 7{,}888 student responses), we analyze whether the difficulty ratings assigned by an LLM match empirical item difficulty. Our study shows that student engagement with LLMs varied across courses and increased over the term. Although students expressed high satisfaction and reported saving considerable time, their perception of deep learning benefits declined, and many noted a tendency toward over-reliance. Regarding the difficulty ratings of LLM-generated MCQs, the Easy, Medium, and Hard labels correlated closely with its assigned Bloom's Taxonomy levels (Spearman $\rho=0.90$), reflecting an artifact of co-generation. However, neither metric predicted empirical item difficulty (difficulty label $\rho=0.06$; Bloom level $\rho=0.02$). The ratings reflect the structural formatting of a question rather than its underlying difficulty.
arXiv ID: 2609.27063 / 要約の誤りについて