コードの動作を変える曖昧さだけを質問するCONTRA
CONTRA: Discovering and Qualifying Behavior-Changing Questions for Selective Clarification in LLM Code Generation
この論文をやさしく読む
ひとことで言うと
曖昧な要件に対する回答を変えたとき、実際にコードの動作が変わるかを調べて、尋ねるべき質問を選ぶ手法です。
何に役立つ?
コード作成前の確認漏れと不要な質問を減らす用途が考えられます。要旨で報告される実証結果は、ベンチマーク上の質問選択性能とプラグイン実装です。
この研究の面白いところ
質問文の意味だけで判定せず、二つの回答に対応するプログラムを共通入力で実行して違いを確認します。質問を続けるかどうかも履歴から決めます。
どこまで分かった?
4エージェントでのF1改善が報告されていますが、日常開発での修正費用や作業時間の削減量は要旨にはありません。比較対象との結果は同一LLM・同一評価手順という条件に基づきます。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
コーディングエージェントは、正しく見えるにもかかわらず、ユーザーが意図していない動作を実装したコードを生成することがある。この不一致は、指定が不十分な要件をエージェントが独自の仮定で黙って補うことで生じうる。その仮定を基に後続の開発が進むと、出来上がった動作の修正は次第に高コストになりうる。早期の確認は不一致の防止に役立つが、不要な質問は開発者の作業を中断し、開発を遅らせる。既存手法は、不要な質問を避けながら重要な確認質問を見つけることが難しい。 そこで本研究では、幅広い質問候補の発見と、意味および実行に基づく質問の適格性判定を組み合わせた、学習不要の手法CONTRAを提案する。まず質問候補を生成し、必要な動作に無関係な質問と、要件内ですでに解決されている質問を除く。残った各質問について、ありうる二つの回答を条件にそれぞれプログラムを生成し、共通の入力に対して安定した動作の違いが生じるかを確認する。その後、やり取りの履歴を使って、適格と判定した質問から次に尋ねるものを選ぶか、質問を終了する。 ClarifyCodeBenchの実験では、CONTRAは4種類すべてのコーディングエージェントで最高のF1を達成し、最良のベースラインのマクロ平均F1を13.88パーセントポイント上回った。同一のLLMと評価手順を使った比較でも、Claude CodeおよびOpenHandsというコーディング用実行基盤より高い確認質問の再現率とF1を達成した。実用を支援するため、日常の開発に選択的な確認を組み込むClaude CodeプラグインとしてもCONTRAを実装した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Coding agents can generate code that appears correct but implements behavior the user never intended. This mismatch can arise when an agent silently resolves underspecified requirements through its own assumptions. As subsequent development builds on these assumptions, correcting the resulting behavior can become increasingly costly. Early clarification can help prevent such mismatches, but unnecessary questions can interrupt developers and slow down development. Existing methods struggle to identify key clarification questions while avoiding unnecessary ones. Therefore, we propose CONTRA, a training-free method that combines broad question discovery with semantic and execution-based question qualification. CONTRA first generates candidate questions and filters out those unrelated to required behavior or already resolved by the requirement. For each remaining question, it generates programs conditioned on two plausible answers and checks for stable behavioral differences on shared inputs. It then uses the interaction history to select among qualified questions or stop asking. Experiments on ClarifyCodeBench show that CONTRA achieves the highest F1 with all four coding agents, exceeding the best baseline macro-average F1 by 13.88 percentage points. With the same LLM and evaluation protocol, CONTRA also achieves higher clarification recall and F1 than the coding harnesses Claude Code and OpenHands. To support practical use, we also implement CONTRA as a Claude Code plugin that integrates selective clarification into everyday development.
著者のコメント
15 pages. Code: https://github.com/fangz-cs/Contra
arXiv ID: 2610.01769 / 要約の誤りについて