arXiv論文メモ
新着一覧
cs.DL / cs.AI / cs.HC · 査読状況未確認

論文の著者が内容を理解しているかを測るgreCAPTCHA

greCAPTCHA: Assessing Understanding as Evidence of Research Authorship Under Generative AI

Justin Payan, Bálint Gyevnár, Atoosa Kasirzadeh, Nihar B. Shah

この論文をやさしく読む

ひとことで言うと

論文の著者が自分の寄与内容を理解し、検証できるかを質問で確かめる評価方式を提案します。

何に役立つ?

投稿原稿への名前だけでは判断しにくい、内容の理解や関与を評価する補助的な仕組みを目指します。

この研究の面白いところ

理解の深さに応じた質問を作り、31人の研究者の試験と面接で評価しました。参加者の著作かどうかを予測する自動得点のAUCは0.90でした。

どこまで分かった?

監督付き条件での初期的な研究です。参加者から導入前の改善も求められており、理解度の得点だけで著者資格や不正を確定する仕組みではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

学会、学術誌、研究助成機関、学校、大学は、人間の著者名で提出されながら、その著者が原稿を十分に監督していない可能性のある、AI生成と疑われる投稿の急増に対応を迫られている。そのため投稿を評価する機関は、提出物に著者名が記載されていることだけを根拠に、その人の専門性を確実に認めることができなくなっている。 この問題に対処するため、研究原稿についての著者の理解を測定する、監督付き評価手法greCAPTCHAを提案する。測定の基礎となる概念は「検証する能力」であり、原稿への自分の貢献を支える内容を批判的に評価するために必要な知識と推論能力と定義する。greCAPTCHAは複数段階の理解を評価する質問を生成し、著者の回答に基づく評価報告を提供する。 試作システムを用い、31人の研究者を対象に利用者調査と半構造化インタビューを行った。自動採点は、論文を調査参加者が執筆したか否かをAUC 0.90で予測した。参加者は全体として肯定的な利用体験を報告し、著者の理解を測るうえで構成概念妥当性が適切だと述べた一方、導入前に必要な重要な変更も提案した。本結果は、監督付きの条件下でgreCAPTCHAが個々の原稿に関する理解を評価できるという初期的な証拠を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Conferences, journals, funders, schools, and universities are struggling with a surge of potentially AI-generated submissions from ostensibly human authors, who may not have exercised sufficient human oversight for their manuscripts. In turn, institutions evaluating submissions can no longer reliably credit expertise based solely on authors' names on submitted work. To address this problem, we propose greCAPTCHA, a proctored assessment approach that measures authors' understanding of research manuscripts via the construct of capacity to verify, which we define as the knowledge and reasoning required to critically assess the contents underlying one's contributions to a manuscript. greCAPTCHA generates questions assessing multiple levels of understanding and provides an evaluative report based on authors' responses. Using a prototype implementation, we conduct a user study and semi-structured interviews with $31$ researchers to evaluate greCAPTCHA. Its automated scores predict which papers were or were not authored by study participants with an AUC of $0.90$. Participants reported positive overall experiences with the system and remarked on the appropriate construct validity for author understanding, while also suggesting important changes to be made before deployment. Our results provide initial evidence that greCAPTCHA can assess manuscript-specific understanding under proctored conditions.

arXiv ID: 2609.20481 / 要約の誤りについて