秘密データを明かさずAI出力を検証できる条件
Can AI Oversight Be Zero Knowledge?
この論文をやさしく読む
ひとことで言うと
AIが秘密情報を使って出した答えを、秘密を見せずに正しいと確認できるかを調べた理論研究です。外部の判断や実験などを参照する計算を対象にします。
何に役立つ?
機密性を保つAI監督の仕組みを設計するとき、どのような追加条件が必要かを考える基礎になります。外部情報源が回答に署名を付ける場合には、ゼロ知識検証が可能になると示しています。
この研究の面白いところ
速い検証ではなく、秘密を漏らさない検証に焦点を移しています。一般の場合の不可能性と、署名付きオラクルでの可能性を同じ研究で対比しています。
どこまで分かった?
不可能性はランダムオラクルモデルでの一般的な計算についての結果です。肯定的結果には署名付き回答と衝突耐性ハッシュ関数の仮定があり、署名だけで人間の判断や物理実験の内容自体が正しくなるという意味ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
AIシステムは、医療記録に基づく職務適性の評価や、秘密の構造に基づく薬剤候補の性質の予測など、機密データから出力を生成することが増えている。基となるデータを明かさずに、こうした出力が正しいことを検証するのは重要である。最近の一連の研究は、人間の判断、物理実験、ウェブなどのオラクルに正しさが依存し得る、オラクル支援計算について、対話的証明や討論を通じたAI出力の検証を研究している。これらは、計算そのものよりはるかに速く動く検証者による検証に注目する。しかし、このような効率的検証は一般のオラクル支援計算では不可能であるため、追加の仮定に依存している。 本研究は代わりにプライバシーに焦点を当てる。検証者が計算量に対して多項式時間で動作することを認めたうえで、オラクル支援計算の対話的論証をゼロ知識にできるか、すなわち、検証者が出力の正しさ以外に機密データについて何も学ばないようにできるかを問う。一般には不可能であることを証明する。ランダムオラクルモデルでは、証明者と検証者の双方が計算自体よりはるかに長い時間動作することを認めても、全てのオラクル支援計算に対するゼロ知識証明は存在しない。この不可能性は、拡張可能な監督の代表的なモデルである討論にも及ぶ。 肯定的な結果として、オラクルが各回答に暗号署名を付けるなら、衝突耐性ハッシュ関数だけを仮定して、効率的な証明者と検証者により、全てのオラクル支援計算をゼロ知識で検証できることを示す。プライバシーに加え、これは拡張可能な監督への別の方法も与える。この方法は、討論での誠実な相手にも、従来の単一証明者プロトコルでの計算の頑健性にも依存しない。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studies verification of AI outputs via interactive proofs and debate for oracle-aided computation, where correctness may depend on an oracle such as human judgment, a physical experiment, or the web. These works focus on verification by a verifier that runs much faster than the computation. However, such efficient verification is impossible for general oracle-aided computation, and these works therefore rely on additional assumptions. We focus instead on privacy: allowing the verifier to run in time polynomial in the computation, we ask whether interactive arguments for oracle-aided computation can be zero knowledge, so that the verifier learns nothing about the confidential data beyond the correctness of the output. We prove that, in general, they cannot. In the random oracle model, there are no zero-knowledge proofs for all oracle-aided computations, even if both the prover and the verifier are allowed to run much longer than the computation itself. The impossibility extends to debate, a canonical model for scalable oversight. On the positive side, we show that if the oracle attaches a cryptographic signature to each of its answers, then every oracle-aided computation can be verified in zero knowledge with an efficient prover and verifier, assuming only collision-resistant hash functions. Beyond privacy, this also gives an alternative approach to scalable oversight that relies neither on an honest opponent, as in debate, nor on the robustness of the computation, as in prior single-prover protocols.
arXiv ID: 2610.01995 / 要約の誤りについて