利用者の圧力下で業務AIが規則を守るか測るPACT
PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
この論文をやさしく読む
ひとことで言うと
業務AIに、急いでいる、例外にしてほしいなどの圧力がかかったとき、指定された規則を守れるかを会話形式で測っています。
何に役立つ?
企業向けAIのモデル選択や、規則違反を防ぐ仕組みの評価に使うことが想定されています。規則を守るだけでなく、適用すべき場面を見分ける能力も評価します。
この研究の面白いところ
12領域48シナリオで複数の圧力と言い回しを使い、単一の正解率では捉えにくい順守の側面を6指標で整理しています。
どこまで分かった?
結果は作成したベンチマークと22モデルでの評価であり、法的適合性の認定ではありません。平均65%は違反率の相対的な増加で、65ポイントの増加ではありません。評価項目の作成監査にもLLMを用いています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
企業でのAI導入が拡大する中、企業向けLLMエージェントは採用、医療、金融などの機微な状況に導入されている。こうした状況では、エージェントのシステムコンテキストに記載された規則を守ることが、最重要の法的懸念となる。現時点では、特に執拗な利用者、急ぐ管理者、あるいは違反が便利で魅力的な状況からの圧力の下で、どのLLMモデルが規則に違反しやすいかを系統的に測る評価枠組みはない。 本研究では、規制のある12の企業業務領域にまたがる日常業務を支援するAIエージェントについて、圧力下での規則順守を測るベンチマークPACT(Pressure-Applied Compliance Testing)を導入する。48のシナリオを、それぞれ現実的な複数ターンの会話として設定する。各評価項目では、常設の規則とそれに違反する近道を対にし、異なる言い回しとシステムプロンプトのモードにわたって、一連の圧力を加える。サンプルが曖昧でなく、評価の抜け道を利用できず、評価されていると察した挙動を引き出さない程度に現実的であるよう、厳格なLLM評価者による監査の下で、PACTを構成要素ごとに作成する。 6つの相補的な指標によってLLMの規則順守を評価し、圧力下および複数ターンの会話を通じた頑健性、透明性、規則が適用される場面を正しく見分ける能力を総合的に捉える。この評価を、全項目と全モードにわたる信頼性で重み付けした順守率PACTScoreに集約する。複数の提供元と規模にまたがる一般的な22のLLMモデルの結果は、モデル間でも指標の次元間でも順守に大きなばらつきがあることを示した。最も強いアシスタントでも6~10%の項目で規則の適用を誤り、通常の利用者からの圧力によって違反率は平均65%増加する。PACTはLLMアシスタントの規則順守上のリスクを示し、安全策と慎重なモデル選択の必要性を提起する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
As corporate AI adoption continues to grow, enterprise-grade LLM agents are being deployed into sensitive contexts such as hiring, healthcare, and finance. In these contexts, compliance with rules specified in an agent's system context is a first-order legal concern. Currently, no evaluation framework systematically measures which LLM models tend to violate compliance rules, especially under pressure from a persistent user, a hurried manager, or circumstances where violation is convenient or attractive. We introduce PACT (Pressure-Applied Compliance Testing), a benchmark for rule-following under pressure in AI agents assisting employees in daily tasks across twelve regulated enterprise domains and forty-eight scenarios, each set in a realistic multi-turn conversation. Each benchmark item pairs a standing rule against a rule-violating shortcut, and applies a battery of pressures across different wordings and system-prompt modes. We construct PACT component by component under strict LLM-as-judge auditing to ensure samples are unambiguous, ungameable, and realistic enough to avoid eliciting evaluation-aware behavior. We use PACT to profile LLM compliance across six complementary metrics that create a holistic picture of an AI assistant's robustness under pressure and throughout multi-turn conversations, its transparency, and ability to correctly discern where a rule applies. We aggregate this profile into PACTScore, a reliability-weighted compliance rate over all items and modes. Our results across 22 common LLM models spanning multiple providers and sizes show substantial variability in compliance across models and metric dimensions. Even the strongest assistants mis-apply a rule on 6 to 10% of items, and ordinary user pressure raises the violation rate by 65% on average. PACT highlights compliance risks in LLM assistants, motivating guardrails and careful model selection.
著者のコメント
26 pages, 12 figures, 17 tables. Includes technical appendix; Dataset: https://huggingface.co/datasets/trace-ai-labs/pact; Code: https://github.com/trace-ai-labs/pact
arXiv ID: 2609.18605 / 要約の誤りについて