arXiv論文メモ
新着一覧
cs.SE / cs.HC · 査読状況未確認

研究者はAIが書いたコードをどう確かめているか

How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming

Gabrielle O'Brien, Reed Milewicz, Nasir Eisty

この論文をやさしく読む

ひとことで言うと

研究者のAIコーディング利用の実例を調べ、コードが動くかの確認は多い一方、自動テストや他者のレビューは少ないことを報告しています。

何に役立つ?

科学計算でAIコードを使う際に、課題に合った検証を支える仕組みを設計する手掛かりになります。

この研究の面白いところ

検証への自信と実際に報告された検証方法が結び付かず、ツールや自分への信頼の方が強く関連していた点です。経験による違いも方法より信頼の向きに現れます。

どこまで分かった?

主に米国の大学の研究者による、各人一つの利用事例の自己報告です。コードの正しさを直接測定した調査でも、検証方法の効果を無作為実験で比較した研究でもありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成AIは研究用のプログラミングに入り込んでいるが、研究者がどの課題を任せ、そのコードが正しいかをどう判断しているかについての証拠は乏しい。本研究では、コードを書く研究者を対象に2025年に行った調査の自由記述527件を用いる。回答者の大半は米国の大学に所属する。各回答では、自身の研究の一つの課題、そのためのAIツールの使い方、結果を評価するために行ったことが述べられている。報告された課題と評価戦略を符号化し、プログラミング経験、研究分野、確信度の評価との関係を調べた。 利用は、データ処理、可視化、デバッグ、数学的・科学的計算、統計解析の5課題に集中していた。評価は非公式で個人に委ねられていた。半数を超える記述が生成コードを実行したと述べる一方、自動テストや他者によるレビューはまれだった。利用事例と評価戦略はプログラミング経験によってあまり変わらなかったが、確信は異なり、経験の少ない人は自分よりAIを信頼し、経験のある人はその逆だった。評価への確信は報告された戦略と関連せず、最も強く関連していたのはツールと自分自身への信頼だった。科学的コードへのAIの寄与の検証は、共有のテスト・レビュー基盤の外で、主に個人の判断に依存していた。インターフェースは、評価を利用者に任せきりにするのではなく、課題に適した評価を支援できる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative AI has entered research programming, yet there is little evidence about which tasks researchers hand to it or how they decide whether its code is correct. We draw on 527 free-text responses to a 2025 survey of researchers who write code, most of them at U.S. universities. In each response, a researcher recounts a single task from their own work, the way they used an AI tool for it, and what they did to assess the result. We coded the task and the evaluation strategies reported, and related both to programming experience, research area, and confidence ratings. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis. Evaluation was informal and individual. Over half of accounts described running the generated code, while automated tests and review by another person were rare. Use cases and evaluation strategies varied little with programming experience, but confidence did: Less experienced programmers trusted the AI more than themselves, and experienced programmers the reverse. Evaluation confidence was not associated with the strategies reported. Its strongest correlates were confidence in the tool and in oneself. Validating AI contributions to scientific code rested largely on individual judgment, outside shared infrastructure for testing or review. Interfaces could support task-appropriate evaluation rather than leave it to the user.

arXiv ID: 2609.22049 / 要約の誤りについて