重大な安全性低下の割合を制約するLLM追加学習
Beyond Average Safety: Chance-Constrained LLM Fine-tuning
この論文をやさしく読む
ひとことで言うと
追加学習で一部の危険な入力への応答だけが大きく悪化する事態を抑える方法です。
何に役立つ?
考えられる用途は、安全性を保ちながらLLMを特定の仕事に追加学習することです。要旨では三つの課題と三つのモデルで既存手法と比較しています。
この研究の面白いところ
平均的な安全性ではなく、所定の閾値を超えて悪化する事例の割合を直接制限します。不連続な条件を微分可能な保守的制約に置き換えています。
どこまで分かった?
実験は要旨に記載された有害な追加学習の三課題と三モデルに関するものです。ほかのモデルや運用条件での保証について、要旨には結果がありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルを新たな目的に合わせて追加学習すると、有用性、指示への追従、特定分野での性能が向上する一方、安全上重要な入力への応答が悪化することもある。既存の安全性を保つ追加学習法は、通常、安全性損失の平均を制御するか、重み付きの補助的な罰則を使うため、まれでも深刻な失敗を見えにくくする。本研究では、参照モデルに比べた安全性の低下が所定の閾値を超える事例の割合を制限する、確率制約付きの定式化を提案する。 この経験的な確率制約には不連続な指示関数が含まれるため、違反率を微分可能な関数で上から抑え、扱いやすい保守的な制約にする。次に、この制約をパラメータ空間における安全な集合として扱い、実行可能性を保つために追加学習の更新方向を最小限だけ修正する、制約を考慮した勾配降下法を開発する。更新には閉じた形の解があり、低下の閾値に近い事例や超えた事例を重視する、安全性低下の裾に着目した補正をもたらす。三つの異なる課題と三つのモデルを使った有害な追加学習に関する広範な実験で、この方法は文献にある比較手法を一貫して上回った。これらの結果は、LLMの追加学習で安全性を保つ問題を、平均リスクへの正則化よりも信頼性制約付き最適化として捉えることを示唆する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Fine-tuning large language models on new objectives can improve helpfulness, instruction following, or domain-specific performance, but it can also induce regressions on safety-critical prompts. Existing safety-preserving fine-tuning methods typically control average safety loss or use weighted auxiliary penalties, which can obscure rare but severe failures. We propose a chance-constrained formulation for safety-preserving fine-tuning that limits the fraction of safety examples whose degradation relative to a reference model exceeds a prescribed threshold. Because the resulting empirical chance constraint contains a discontinuous indicator, we introduce a differentiable majorization of the violation rate, yielding a tractable conservative constraint. We then develop a constraint-aware gradient descent method that treats the majorized constraint as a safe set in parameter space and minimally modifies the fine-tuning direction to preserve feasibility. The resulting update admits a closed form and produces a tail-aware safety correction that emphasizes examples near or above the degradation threshold. We conduct an extensive set of experiments on harmful fine-tuning across three different tasks and three models and show that our approach consistently outperforms the baselines that exist in the literature. These results suggest that safety preservation in LLM fine-tuning is better viewed as a reliability-constrained optimization problem than as average-risk regularization.
arXiv ID: 2609.29960 / 要約の誤りについて