第三者がAIモデルの同一性を確認する質問応答方式
TP-CRIV: A Framework for Third-Party Challenge-Response Identity Verification of AI Models
この論文をやさしく読む
ひとことで言うと
公開サービスのAIモデルと、所有を主張する側が持つモデルの同一性を、第三者が新しい課題への応答から統計的に調べる方式。
何に役立つ?
遠隔提供されるモデルの流用を疑う場面で、提供者の特別な協力や主張者のモデルへの直接アクセスなしに証拠を得る用途が考えられる。
この研究の面白いところ
未公開の課題とネットワーク隔離を組み合わせ、課題を見てから外部の助けで答えることを防ぐ設計にしている。
どこまで分かった?
証拠は統計的であり暗号学的な証明ではない。要旨での具体的な実験対象は、ImageNetで事前学習したTorchVisionのCNN画像分類モデル10件である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
AIモデルが遠隔サービスを通じて提供されることが増え、モデルの不正流用への懸念が高まっている。透かし、指紋、モデルの類似性解析など既存の方法は主に、事前に定めた証拠や挙動の直接比較に依存し、同一性を主張する当事者が、そのモデルに関する情報を現在保持して利用できるかを明示的には評価しない。本論文はAIモデルのための第三者によるチャレンジ・レスポンス式同一性検証、TP-CRIVを提案する。 TP-CRIVが対象とするのは、検証者が主張者のモデルを内部から調べることもAPIで利用することもできず、疑わしい公開サービスとは通常のブラックボックス推論インターフェースを通じてしかやり取りできず、サービス提供者にこの検証手順への特別な協力を求めない状況である。この制約下で、主張者が公開モデルに対して事前に定めた同一性を満たすモデルを手元に保持しているかどうかについて、経験的な証拠を得られるようにする。検証は新しく未公開の要件とネットワーク隔離の下で行い、課題を明かした後のオンラインの外部支援に依存して能力を示すことができないようにする。得られる証拠は、独立に定めて較正した一致・不一致の運用状況に対して解釈するもので、暗号学的な証明ではなく統計的な証拠である。 CNN画像分類器については、確率制御に基づく証拠生成によってTP-CRIVを具体化した。ImageNetで事前学習したTorchVisionの10モデルを用いた実験では、同一モデルと異なるモデルを明確に分離し、独立に較正した閾値で有限回の課題による検証を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Artificial intelligence (AI) models are increasingly deployed through remote services, making model misappropriation a growing concern. Existing approaches, including watermarking, fingerprinting, and model similarity analysis, primarily rely on predefined evidence or direct behavioral comparison and do not explicitly evaluate whether the claimant currently possesses and can utilize model-dependent information relevant to the claimed model identity. In this paper, we propose Third-Party Challenge-Response Identity Verification (TP-CRIV) for AI models. TP-CRIV targets a third-party verification setting in which the verifier has neither white-box nor API access to the claimant's model, can interact with the suspicious deployed service only through its ordinary black-box inference interface, and does not require protocol-specific cooperation from the service provider. Under these constraints, the framework enables the verifier to obtain empirical evidence as to whether the claimant locally possesses a model satisfying a predeclared identity relative to the deployed model. Verification is conducted under fresh, previously undisclosed requirements and network isolation, so that the demonstrated capability cannot rely on online external assistance after challenge disclosure. The resulting evidence is interpreted relative to independently specified and calibrated matching and non-matching operating situations and is statistical rather than cryptographic. We instantiate TP-CRIV for CNN image classifiers using probability-control-based witness generation. Experiments on ten ImageNet-pretrained TorchVision models demonstrate clear same/cross-model separation and finite-challenge verification using independently calibrated thresholds.
arXiv ID: 2609.29264 / 要約の誤りについて