arXiv論文メモ
新着一覧
physics.ed-ph · 査読状況未確認

選択問題の誤答の偏りから共通の誤概念を見つける

Determining the degree of randomness in multiple-choice question distractors

JPW Diener, BL Frick, and J Kriek

この論文をやさしく読む

ひとことで言うと

誤答が一つの選択肢に集中しているか、ばらばらかを測り、クラスで共有される誤解を詳しく調べるための手掛かりにします。

何に役立つ?

授業後の選択問題から、再説明が必要そうな項目を素早く絞るのに役立ちます。詳しい面接や分析をどこに使うか決めるための予備的な方法です。

この研究の面白いところ

誤答のばらつきを捨てずに測定対象とし、正答率だけでは区別しにくい回答パターンを扱います。独立した面接で確認された誤概念との対応も調べています。

どこまで分かった?

要旨は標本サイズに依存しない枠組みと述べますが、どんな人数でも診断精度が保証されるという数値的結果は示していません。著者も詳細な診断を置き換えるものではなく、一度の実施で行う予備的な振り分けとして位置付けています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

多肢選択問題は効率がよいため教育評価で広く使われるが、その診断能力には本質的な限界がある。従来の心理測定手法は、あらかじめ定められたデータのパターンにモデルを当てはめて誤概念を推定するため、そうした事前の傾向が存在しない、小規模、新規、または標準化されていない標本の分析には有効でない。本研究ではこの問題に対し、情報理論に基づく新しい誤答選択肢分析法を導入する。誤答に含まれるランダムさを雑音としてではなく、学生の回答パターンに現れる測定可能な信号として扱う。 誤答選択肢のエントロピーを全体の正答率と併せて調べることで、標本サイズに依存しない成績分類の枠組みを提供する。標準的な力学の概念テストについて、得点の低い項目の回答データを、独立に記録された誤概念と照合した。その結果、最も集中している、すなわちランダムさの度合いが最も低いと判定された項目では、誤答の過半数が、その項目に対応付けられ、面接でも検証された特定の誤概念を表す選択肢を選んでいた。さらに、ほかの研究で診断上の信頼性が低いと独立に指摘された項目は、この指標がランダムさの度合いを最も高いと判定する項目と正確に一致した。 教育上の具体的な対応につながるのは、クラス内の誤答のランダムさが低い場合であり、これは共通して存在し、再指導の対象にできる誤概念を示す。一方、ランダムさが高い場合は、主として、その集中度を判断するための帰無的な比較例として機能する。本分析は、一度の実施で素早く行う予備的な振り分けとして意図されている。すなわち、従来の誤答選択肢分析法を補完する、実務的で理論に基づいた手法であり、どの項目を、より資源をかけて詳しく調べるべきかを教員が見極める助けとなる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Multiple-choice questions are a staple of educational assessment due to their efficiency, but their diagnostic ability is inherently limited. Traditional psychometric techniques infer misconceptions by fitting models to predetermined data patterns, making them ineffective for analysing small, novel, or non-standardised samples where such prior trends are absent. To address this, we introduce a new distractor analysis method based on Information Theory, which treats the randomness in incorrect answers not as noise but as a measurable signal in student answer patterns. By examining the entropy of distractor choices alongside overall accuracy, our approach offers a sample-sized-independent framework for classifying performance. Cross-referencing low-scoring item-level response data against independently documented misconceptions on a standard mechanics inventory shows that, for items identified as most concentrated (lowest degree of randomness), a majority of incorrect responses select the exact response coded to that item's named, interview-validated misconception. Furthermore, items independently flagged elsewhere as diagnostically unreliable are exactly the items the measure identifies as having the highest degree of randomness. The pedagogically actionable cases are those with low degree of randomness in a class's incorrect answers, signalling a shared, re-teachable misconception, whereas a high degree of randomness chiefly serves as the null case against which that concentration is judged. Our analysis is intended as a fast, single-administration triage step: a practical, theory-driven complement to established distractor-analysis methods that helps instructors flag which items warrant closer, more resource-intensive investigation.

arXiv ID: 2610.01431 / 要約の誤りについて