arXiv論文メモ
新着一覧
cs.CR / cs.AI · 査読状況未確認

安全性でAIを振り分ける評価は分布変化で偏る

False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift

Amit Singh Bhatti, Vishal Vaddina

この論文をやさしく読む

ひとことで言うと

安全なAIを選ぶ振り分け機能について、比較モデルの選び方だけで評価が偏る問題と、攻撃の認識を自己申告させる防御の弱点を調べています。

何に役立つ?

ルーターの改善を正しく測るため、評価用データと比較モデルの選定を分ける設計に役立ちます。攻撃を認識したかだけでなく、実際の有害性指標を測る必要性も示します。

この研究の面白いところ

比較基準の選択による差が、ルーティングの性能不足と見なされた差全体に匹敵しました。また認識率を下げる攻撃が、必ずしも有害性の増加に直結しないことも示しています。

どこまで分かった?

結論は評価した7コーパスや除外分割などの条件に基づきます。制御器に攻撃を組み込んだ結果はオフラインの反実仮想推定で、実運用で観測した被害量ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

安全性ルーターは、各依頼を複数のモデルのいずれかへ送り、最良の単一モデルを基準に評価される。主要なルーティングの基準評価では、比較対象のモデルを評価データ上で選んでいる。その基準評価自身の設定では問題がないが、分布が変わると問題になる。HELM Safetyでは、この選択によるコストは、ランダム分割で有害性指標の0.003〜0.030、カテゴリを学習側から除いた評価で0.045〜0.113であり、ルーティングに帰せられる性能差全体と同程度となる。この方向性は、公表された2つの評価者のどちらを単独で使っても保たれる。AgentDojoでは、スイートを除外して評価すると7〜9倍に増える。 事前に固定した規則で選んだ7つの安全性コーパスのうち、3つが事前登録した区間検定を満たし、4つが後に設定した置換による帰無基準を上回った。区間検定を満たさなかった4つのうち3つは、一部のモデルで観測された有害性がゼロのコーパスである。この偏りの方向は先行研究で証明されている。本研究では有害性と正解率について大きさを測り、今回測定した除外分割でより大きくなることを示し、楽観的な偏りと分布変化に依存するリグレットの和で上から抑える。 分布変化のもとで適切に採点すると、これらの基準評価でルーティングがもたらす利点は小さい。モデルプールの評価セルの大半で、入れ子構成のルーターは、適切に選んだ基準モデルと同じモデルを使う。ほぼ性能が飽和したAgentDojoコーパスでは、振り分け前に完全な判断ができるルーターでも、有害性に関する利得は最大2ポイントである。 また、途中の後段で現れる注入を認識したというモデルの表明は、誘導可能であることも見いだした。未使用の条件での再実行では、相手モデルを知る攻撃者がGPT-5.4の評価上の認識率を19.6ポイント下げ、独立したラベルでも確認された。この攻撃を制御器へオフラインで反実仮想的に組み込むと、代替先のモデルによって推定有害性が上がる場合も下がる場合もあった。安全性ルーティングは、テストラベルを使わずに選んだ基準モデルと比べ、分布変化のもとで評価すべきである。また、認識に基づく防御は、モデルが見る内容を選ぶ攻撃者に対して、有害性そのもので評価すべきである。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its direction holding under either published judge alone. It rises seven- to ninefold on AgentDojo when suites are held out. Across seven safety corpora chosen by rules fixed in advance, three meet a registered interval test and four beat a later permutation null, and three of the four interval misses are corpora where some models have zero observed harm. Prior work proves the direction of this bias. We size it on harm and accuracy, show that it is larger under the held-out splits we measure, and bound it by optimism plus a shift-dependent regret. Scored honestly under shift, routing buys little on these benchmarks. In most pool cells the nested router serves the honest baseline's model, and on the nearly saturated AgentDojo corpus a perfect pre-dispatch router is worth at most two points of harm. We also find a model's expressed recognition of a late injection steerable. On held-out reruns an attacker who knows which model it faces lowers GPT-5.4's judged recognition by 19.6 points, confirmed by an independent label. In an offline counterfactual composition into a controller, the same attack raises or lowers estimated harm depending on the fallback model. Safety routing should be evaluated under shift, against a baseline chosen without the test labels, and recognition-based defences should be scored on harm against an attacker who chooses what the model sees.

arXiv ID: 2610.01535 / 要約の誤りについて