arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

学生の文章へのLLMの助言は教師の着眼点に合うか

Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing

Norah Almousa, Shayan Peyghambari Oskoui, Raquel Coelho, Gayle Rogers, Xiang Lorraine Li, Diane Litman

この論文をやさしく読む

ひとことで言うと

学生の文章にAIが出す助言を教師の助言と比べ、何を指摘するか、原稿段階や成績に合わせて変えるかを調べた。

何に役立つ?

文章指導向けLLMを、単に多くの項目を指摘するかではなく、教師に近い着眼点と適応性で評価する基準になる。

この研究の面白いところ

七種類の着眼点を使い、教師、六つのLLM、三種類のプロンプト方針を同じ枠組みで比較した。

どこまで分かった?

大学の三つの文章作成科目での助言に関する比較である。助言によって学生の学習成果がどう変わったかは、この要旨の結果には含まれない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

最新の大規模言語モデル(LLM)が学生の文章に生成する助言について、着眼点と適応性の面で熟練教師の教育実践を反映しているかを調べる。従来の評価は助言の特徴、学習への影響、対象を調べてきたが、何に焦点を当てるか、状況にどう合わせるかは十分に検討されてこなかった。 この不足を補うため、Narcissの分類を採用・改訂し、助言の着眼点を七種類に分けて、大学の三つの文章作成科目における教師とLLMの助言に注釈を付ける。六つのLLMと三種類のプロンプト方針による助言、教師の助言に注釈を付けたベンチマークFeedTypeを公開する。着眼点の種類をどれだけ網羅するか、その分布はどうかを評価し、原稿の段階や学生の成績水準に応じて、熟練教師のように助言を変えるかを調べる。 ほとんどのLLMは着眼点の種類の大半を扱えたが、教師の助言における分布を再現できず、適応の程度もモデルごとに異なった。教師の適応的な振る舞いに一致したモデルはなかった。著者らは、FeedTypeが今後のLLMによる助言と教育的実践の一致を調べる研究を支えると考えている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined feedback characteristics, its impact on learning, and its target, yet the focus of feedback and its adaptivity remains largely overlooked. To bridge this gap, we adopt and refine Narciss's taxonomy into seven feedback focus types to annotate teacher and LLM-generated feedback across three university writing courses. We release FeedType, a benchmark containing annotated teacher and LLM feedback from six LLMs under three prompting strategies. We assess the coverage and distribution of feedback focus types, and examine whether LLMs adapt their feedback across draft stages and student performance levels as an expert instructor does. Our findings show that while most LLMs cover most feedback focus types, they fail to reflect teacher feedback distributions and show varying levels of adaptivity, with none matching the teachers' adaptive behavior. We believe FeedType will support future research on pedagogical alignment in LLM feedback generation.

著者のコメント

Accepted at AIME-Con 2026. Camera-ready version

arXiv ID: 2609.28026 / 要約の誤りについて