監視の死角を使わない言語モデル部品を秘密のまま証明
ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor
この論文をやさしく読む
ひとことで言うと
追加学習部品が監視器の見えない経路を使わないように設計し、その性質を部品の秘密情報を明かさず証明する方法です。
何に役立つ?
考えられる用途は、第三者製アダプターを受け入れる際の性質の検証です。利用者は配布者への信頼だけに頼らず、宣言された監視器との関係について証明を確認できます。
この研究の面白いところ
悪い部品を統計的に見分けるのではなく、特定の攻撃経路を使えない構造にして証明します。重い解析を公開基盤モデル側に寄せ、非公開部分については一つの線形関係だけを証明する設計です。
どこまで分かった?
保証は宣言された監視器に対する特定クラスの隠れた経路に限られ、すべてのバックドアや有害出力がなくなるという意味ではありません。残差には公開基盤モデル由来の下限があり、適応コストがほぼゼロという結果は5億パラメータのモデルについて述べられています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
重み公開型言語モデル向けの第三者製アダプターは、不透明な重み行列として配布される。受け手は、配布者を信頼するか、配布者の中核資産である重みを調べることなしには、バックドアが隠されているかを確かめられない。重要な一つの種類、すなわち安全監視器が構造的に見えない場所にペイロードを置く攻撃に対しては、検出は防御として確実ではない。宣言された監視器を経由する検出器はいずれも、その不可視部分空間上で不変であり、良性の適応もその部分空間を使うため、本研究で評価した不可視部分空間のすべての統計量で、正常なアダプターとバックドア入りアダプターが重なり合うからである。 この経路を検出する代わりに、構造的に存在しないようにし、それを実現したことを証明する。配布者は、監視器がカバーする方向だけを通して入力を読み取るようアダプターを構築し、認証対象となる読み取り因子について何も明かさず、ゼロ知識でこの性質を証明する。証明書が低コストなのは、計算量の大きい監視器の死角の特定が、公開された基盤モデルの決定論的な関数だからであり、証明するのは一つの線形恒等式だけでよい。提供時に残る残差は、証明者が選ぶ許容値ではなく、基盤モデル自身が持つ公開の下限である。 こうして得られるServeGuardは、サプライチェーンの基本要素である。配布者は証明付きアダプターを配布し、その証明によって利用者や規制当局は、認証対象の読み取り因子を知ることも配布者を信頼することもなく、宣言された監視器に対して、この種類の隠れた経路がアダプターにないことを検証できる。受け入れ時の型検査ガードは、この保証を提供時に受け入れたアダプターのバイト列に結び付ける。 4系列、最大70億パラメータの8チェックポイントで調べたところ、監視予算はアーキテクチャに依存した。測定された限界は、グループ化クエリのチェックポイントではvalue経路のランクで飽和したが、マルチヘッドのものではそうならなかった。5億パラメータのモデルでは、良性の適応に対する制約のコストはほぼゼロであり、安全性を左右する要素は監視器の品質となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Third-party adapters for open-weight language models ship as opaque weight matrices; a recipient cannot check whether an adapter hides a backdoor without trusting the publisher or inspecting the weights, the publisher's core asset. For one important class (payloads placed where a safety monitor is structurally blind), detection is unsound as a defense: every detector that factors through the declared monitor is invariant on its blind subspace, and honest and backdoored adapters overlap on every blind-subspace statistic we evaluate, because benign adaptation uses that subspace too. Rather than detect this channel, we make it structurally \emph{absent} and prove that we did. The publisher builds the adapter to read the input only through directions the monitor covers and proves this in zero knowledge, revealing nothing about the read factor it certifies. The certificate is cheap because the expensive part, identifying the monitor's blind spot, is a deterministic function of the \emph{public} base model, so only one linear identity is proved; the served residual is the base model's own public floor, not a prover-chosen tolerance. The result is \emph{ServeGuard}, a supply-chain primitive: the publisher ships a \emph{proof-carrying adapter} whose proof lets a consumer or regulator verify, without the certified read factor and without trusting the publisher, that the adapter carries no hidden channel of this class relative to the declared monitor; an admission-time typing guard binds the guarantee to the adapter bytes admitted at serving time. Across eight checkpoints up to 7B from four families, the monitoring budget is architectural: the measured frontier saturates at the value-path rank on grouped-query checkpoints but not on multi-head ones. On a 0.5B model confinement is nearly free for benign adaptation, making monitor quality the security lever.
著者のコメント
30 pages, 2 figures, and 4 tables
arXiv ID: 2609.21515 / 要約の誤りについて