arXiv論文メモ
新着一覧
cs.CY · 査読状況未確認

大規模言語モデルによる実践的判断力の記述式評価

Automating Constructive Assessment with Large Language Models: Toward Scalable and Repeated Evaluation of Practical Competence

Satoshi Takahashi, Atsushi Yoshikawa, Megumi Kose, Kenichi Suzuki, Chieko Inoue, Yumi Watanabe, Mari Sawada

この論文をやさしく読む

ひとことで言うと

記述式の事例問題づくり、採点、誤答への説明を大規模言語モデルで自動化し、教育評価として検証した研究。

何に役立つ?

実践的判断力を繰り返し測る評価を設計する際、作問や採点の負担を減らす方法の検討に役立つ。

この研究の面白いところ

追加学習をせずプロンプトだけで三つの工程を扱い、内的一貫性の指標は0.78、一部条件では人間の採点との一致率が100%だった。

どこまで分かった?

採点の完全一致は一部の条件での結果であり、あらゆる科目や回答に当てはまるとは示されていない。要旨には、長期的な教育効果の実証結果はない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本研究は、実践的な判断力を評価する構成的な方法である階層的診断推論(HDR)を、大規模言語モデルを用いて自動化する評価手順を開発・検証した。HDRは、事例に基づく問題の誤りを学生に特定・説明させ、高次の認知能力を測る記述式課題であるが、作問と採点に専門知識と労力を要する。そこで、教育上の意図に沿った誤りを含む事例問題の生成、記述式回答の採点、誤答に基づく構造化フィードバックの作成を、追加学習を行わずプロンプト設計だけで自動化し、実証的に検証した。生成問題の得点分布とクロンバックのα係数0.78は、内的一貫性と構成概念妥当性を支持した。自動採点と人間の採点の一致率は、一部の条件で100%に達した。フィードバックは、人間の指導者によるものと同程度に説得力と有用性があると評価された。著者らは、再現性、即時性、低コストを備えたHDR型評価の実施方法を示したとし、大規模言語モデルの柔軟性により、構造を保った反復・縦断的な評価も可能になると述べている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

This study aimed to automate hierarchical diagnostic reasoning (HDR), a constructive method for evaluating practical judgment skills, by developing and testing an evaluation process using a large language model. HDR is a descriptive task that measures higher-order cognitive skills by requiring students to identify and explain errors in case-study-based problems. However, it requires expertise and effort to develop and evaluate. Hence, we proposed and empirically validated the automatic (1) generation of case problems containing errors aligned with educational intentions, (2) scoring of descriptive answers, and (3) generation of structured feedback based on incorrect answers, achieved solely through prompt design without fine-tuning. The internal consistency and construct validity of the generated problems were supported by the experimental score distribution and Cronbach's alpha (0.78). The agreement between automated and human ratings reached 100% under some conditions. The feedback was rated as being as convincing and useful as that from human instructors, demonstrating a practical framework for implementing HDR-based constructive assessment with reproducibility, immediacy, and low cost. The flexibility of large language models will also enable repeated and longitudinal assessments while maintaining structure, showing broad potential for application in educational settings.

著者のコメント

23 pages, 4 figures, 9 tables

arXiv ID: 2609.25790 / 要約の誤りについて