安全方針を先に宣言して回答との整合性を監査
AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment
この論文をやさしく読む
ひとことで言うと
回答前に安全上の方針を構造化して宣言させ、その方針の正しさを回答への報酬条件にする学習方法です。
何に役立つ?
危険な回答と無害な質問への過剰拒否を減らしつつ、事前方針と回答の整合性を検査するために役立ちます。
この研究の面白いところ
安全計画と回答の評価を単に足し合わせるのでなく、計画の正しさを回答報酬の入口にしています。
どこまで分かった?
改善値は記載されたQwenモデル群の評価結果です。要旨ではASRとLSRの正式な定義や評価データの詳細が示されておらず、あらゆる要求で安全が保証されるという結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
安全性調整のパイプラインは最終回答だけを評価するため、頑健な拒否と、二つの望ましくない近道を区別しにくい。その近道とは、無害な要求まで一律に拒否することと、見栄えはよいが実際には回答を制約しない、不忠実な安全性の理由づけである。本研究は、単一モデルがまず簡潔な構造化安全計画を出力し、その計画を条件として回答する、計画後回答方式AUDITPLANを提案する。計画には脅威ラベル、予定する行動、明示的な制約を記録し、実運用では利用者に表示しないまま、機械的に確認できる監査を可能にする。 この振る舞いを教師ありファインチューニングで学習させた後、安全計画が正しい場合にのみ回答への報酬を与える報酬ゲート目的FAITHGATEを用いて、強化学習を行う。これにより、安全そうに見えるだけで計画に忠実ではない振る舞いを抑え、計画と回答の結びつきを強める。 Qwen系の基盤モデルで、AUDITPLANは頑健性と監査可能性の両方を改善する。Qwen2.5-3B-Instructでは、FAITHGATEによってASRが24.0%から11.6%へ、LSRが1.0%から0.36%へ、過剰拒否が11.0%から2.0%へ低下し、回答だけを対象とする強化学習、自由形式の説明、構造化報酬の重み付き和を上回る。同様の傾向はQwen2.5-1.5B-Instructでも見られる。Qwen-3-4B-InstructおよびQwen2.5-7B-Instructによる、より大きなモデルでの確認実行でも同じ傾向が保たれ、明示的な内部コミットメントが、安全性調整をより忠実で頑健かつ監査可能にすることを示唆する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this behavior with supervised fine-tuning followed by reinforcement learning with FAITHGATE, a reward-gating objective that grants answer reward only when the safety plan is correct. This discourages safe-looking but unfaithful behavior and promotes tighter plan-answer coupling. Across Qwen backbones, AUDITPLAN improves both robustness and auditability: on Qwen2.5-3B-Instruct, FAITHGATE reduces ASR from 24.0% to 11.6%, LSR from 1.0% to 0.36%, and over-refusal from 11.0% to 2.0%, outperforming answer-only RL, free-form explanation, and weighted-sum structured rewards. Similar trends hold for Qwen2.5-1.5B-Instruct. Larger-model confirmation runs on Qwen-3-4B-Instruct and Qwen2.5-7B-Instruct preserve the same trend suggesting that explicit internal commitments can make safety alignment more faithful, robust, and auditable.
arXiv ID: 2609.19325 / 要約の誤りについて