プラグイン更新スキルの評価結果を契約条件まで検証
Evaluating Agent Skills for Version-Specific Plugin Migration: A Retrospective Study
この論文をやさしく読む
ひとことで言うと
プラグイン更新を助けるAI向けスキルの点数が、実際のバージョン要件を満たす助言を示しているか調べました。
何に役立つ?
コード移行を支援するエージェントを評価するとき、総合点だけでなく要件ごとの証拠と採点誤りを確認する手順に役立ちます。
この研究の面白いところ
平均点は上がりましたが、採点を詳しく見直すと改善の信頼区間がゼロに接するかまたぎました。
どこまで分かった?
対象は1つの配布済みスキルと16課題です。実際に動く端から端までの修正や独立した人手評価は今後の課題です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
エージェント向けのスキルは、特定バージョンの保守知識をまとめられる。しかし診断の点数が高いだけでは、移行の助言が対象バージョンの契約を満たすとは言えない。本研究は、配布済みのプラグイン更新スキルを、16件の固定された移行課題に関する64件の報告記録で遡及的に調べた。各条件で2回試行し、評価基準について328件の判定があった。スキルを使うと記録された報酬の平均は93.83から98.75へ上がり、差は4.92ポイント、課題単位のブートストラップによる95%区間は[0.31, 10.86]だった。ただし改善は1課題に集中し、8組の課題では点数が上限に達していた。すべての判定を対応する契約の領域まで追い、10件の報告を詳しく調べると、どちらの条件にも有利になり得る採点誤りが見つかった。ある例では、親ディレクトリを受け入れる包含条件に満点が与えられていた。実行できる検査によってこの欠陥を確認し、動作する終了処理の修正が、より狭いライフサイクル評価基準によってのみ除外されることも示した。精査した判定を置き換えても改善の推定値は4.61~5.39ポイントで正だったが、その区間はゼロに接するかまたいだ。さらに、条件名と元の点数を見せず、別のモデル系列2種類の判定器で64件すべてを再採点した。元の判定器との一致率は91.8%と95.7%で、重み付きκはそれぞれ0.64と0.72、推定改善は10.63と6.09ポイントだった。この研究は、集計した報酬を契約条件ごとの証拠と判定器への感度に結び付ける追跡可能な評価と、移行助言を点検する具体的な方法を示す。実行可能な端から端までの修正、人手による独立した注釈、ほかの枠組みは今後の課題とする。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Agent skills package version-specific maintenance knowledge for coding agents, but a higher diagnostic score does not by itself show that the resulting migration advice satisfies the target version's contract. We study a shipped plugin-upgrade skill through an archive of 64 reports on 16 static migration tasks, with two attempts per condition and 328 criterion decisions. With the skill, mean recorded reward rises from 93.83 to 98.75, a gain of 4.92 points (95% task-bootstrap interval [0.31, 10.86]); the gain is concentrated in one task, and eight task pairs are at the ceiling. Tracing every decision to its contract domain and reviewing ten reports in depth exposes grading errors that favor either arm; in one, a containment predicate that accepts the parent directory still receives full credit. Executable probes confirm this defect and show that a working teardown repair is excluded only by a narrower lifecycle rubric. Replacing the reviewed decisions keeps the estimate positive (4.61 to 5.39 points) but moves its interval to or across zero. Re-grading all 64 reports with judges from two other model families, without arm labels or prior scores, agrees with the original judge on 91.8% and 95.7% of decisions (weighted $\kappa=0.64$ and $0.72$) and gives gains of 10.63 and 6.09 points. The study contributes a traceable evaluation that connects aggregate reward to contract-level evidence and judge sensitivity, together with concrete review checks for migration advice. Executable end-to-end repairs, independent human annotation, and other frameworks are left to future work.
著者のコメント
23 pages, 6 figures, 5 tables. Code, data, and evaluation artifacts: https://github.com/oh-my-dsh/dsh-plugin-upgrade-skill
arXiv ID: 2609.30120 / 要約の誤りについて