SQL生成の保留判断は正解判定法でリスクが変わる
Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL
この論文をやさしく読む
ひとことで言うと
自然言語から生成したSQLを保留する安全性評価が、正解ラベルの作り方で大きく変わると調べた研究です。
何に役立つ?
SQL生成システムの信頼度や保留判断を評価する際に、判定基準を複数使い、専門家監査を組み込む必要性を示します。これは評価方法についての結果であり、モデルの本番環境での安全性を保証するものではありません。
この研究の面白いところ
単一データベースと複数インスタンスのテスト群を入れ替える事前登録実験に加え、専門家の盲検ラベルで両方のずれを確かめています。指標を作った判定基準でそのまま評価すると良く見える問題も検査しています。
どこまで分かった?
対象はSpider-Realistic、SQL特化の4チェックポイント、2つの分割方法です。専門家ラベルの比較は要旨で述べる2チェックポイントについてであり、他の自然言語処理タスクへの一般化は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自然言語からSQLを作るシステムで、不確かな回答を保留するための適合的な保証は、調整に使う正誤ラベルが正確であって初めて信頼できる。実行結果の一致から信頼度を読む不確実性評価の処理は、ベンチマークに付属する単一のデータベースからラベルを得るが、この判定基準は甘いことが知られている。本研究は事前登録した介入実験をSpider-Realisticで行い、そのデータベースをベンチマークの複数インスタンスを蒸留したテスト群に置き換えた。SQLに特化した4つのチェックポイントと2つのデータ分割方法では、置き換え後の保証の未使用データ上のリスクは、自身のラベルが報告するリスクより2.73〜10.23ポイント高くなった。しかし、どちらの判定基準も専門家が評価するリスクを報告できなかった。2人のSQL専門家による盲検のラベルでは、名目上0.10のリスクに調整された保証が、2つのチェックポイントで20.0ポイントと17.2ポイントのリスクを伴った。厳しい方の判定基準も両方向に誤る。却下した回答の大半は専門家から誤りと判定されず、受け入れた回答の中にも誤りがあった。却下回答についてAIで全数調査すると、対象集団によって4分の1から3分の1に意味上の誤りが見つかった。残りの大半は、問題文の指定不足、人工的なデータ例、参照SQLの欠陥の疑いに分類され、最後の懸念は事前登録した専門家の盲検監査でも裏づけられた。判定基準は信頼度の評価結果も左右する。実行結果の一致に基づく信頼度指標は、指標のクラスタを作ったのと同じ判定基準のラベルで評価すると、16通りすべてで良く見えた。専門家ラベルで評価すると、付属データベースではなくテスト群のクラスタで指標を作ることで、ROC曲線下面積(AUROC)が一方のチェックポイントで6.96ポイント、もう一方で1.53ポイント上がった。後者では、専門家評価の区間はテスト群のラベルが報告する8.3ポイントという値を含まなかった。著者らは、保証の結果を両方の判定基準で報告し、基準間の差を意味上のリスクと読む前にベンチマークを監査すべきだとする。また、一致に基づく指標は、その指標を構築したものとは別の判定基準で評価すべきだとする。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
A conformal abstention certificate for text-to-SQL is only as truthful as the correctness labels it is calibrated on. The uncertainty pipelines that read confidence off execution consistency take those labels from the single database a benchmark ships, an oracle known to be lenient. We run a preregistered intervention on Spider-Realistic, swapping that database for the benchmark's distilled multi-instance test suite. Across four SQL-specialist checkpoints and two split schemes, the swap raises the certificate's held-out risk 2.73 to 10.23 points above the risk its own labels report. Neither oracle reports the risk experts assign. Under blinded labels from two SQL experts, a certificate calibrated at a nominal 0.10 carries 20.0 and 17.2 points of risk on two checkpoints. The stricter oracle errs in both directions: most of the answers it rejects are not judged wrong, and some of those it accepts are. An AI-assigned census of what it rejects finds a semantic error in a quarter to a third of them, depending on the population. It attributes most of the rest to underspecified questions, synthetic instances or suspected reference-query defects, a flag supported by a preregistered blinded expert audit. The oracle also decides how a confidence score is judged. Every execution-consistency score looks better under the labels of the oracle that built its clusters, in 16 of 16 combinations. Under expert labels, building such a score on suite clusters instead of shipped-database clusters raises its area under the ROC curve (AUROC) by 6.96 points on one checkpoint and 1.53 on the other. On the second, the expert interval excludes the 8.3 points the suite labels report. A certificate should be reported with both oracles, and an oracle-relative difference read as semantic risk only after the benchmark is audited. A consistency score should be evaluated under an oracle that did not build it.
arXiv ID: 2609.25938 / 要約の誤りについて