arXiv論文メモ
新着一覧
cs.CR / cs.AI · 査読状況未確認

生成AIの行動規則を定義し適用する方法を整理

From Alignment to Access Control: A Framework for GenAI Policy Enforcement

Nathalie Baracaldo

この論文をやさしく読む

ひとことで言うと

生成AIに守らせたい行動規則を、どのように定義し、実際に適用しているかを整理する研究です。

何に役立つ?

開発や運用の担当者が同じポリシーという言葉で異なる仕組みを指していないかを確認し、既存手法を比較するために役立ちます。

この研究の面白いところ

AIの行動を望ましい方向に合わせる考え方からアクセス制御まで、分断されがちな手法を体系的に分析する枠組みを提案しています。

どこまで分かった?

要旨は分析方法と提言を中心としており、特定の実装で事故率が下がったという定量評価は示していません。講演の拡張論文という位置付けから、査読や規則適合の認定を推測することはできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成AIアプリケーションは大きく広がり、利用者が大規模言語モデルと会話したり、さまざまな作業を代行するエージェントを作ったりできるようになった。この分野の能力開発は極めて速く、安全性とセキュリティは後回しになっている。残念ながら、安全性とセキュリティの仕組みの進歩が遅れたことは、実際の事故につながっている。ポリシーはアプリケーションの望ましい振る舞いを定義できるため、システムを安全で規則に適合したものにする基礎である。ところが、ポリシーという言葉の意味は実務者によって異なり、そのことが混乱や、規則への適合には不十分な分断された解決策を生んでいる。 本論文では、生成AIアプリケーションでポリシーを適用する際の、良い点、悪い点、厄介な点を概観する。実際に使われているポリシーの定義・適用手法を体系的に分析し、要素に分解するための方法論を提案する。この原則に基づく分析を踏まえ、提言と、コミュニティに取り組みを求める課題を示す。本論文は、著者Nathalie BaracaldoによるUSENIX Security 2026 Enigma講演「From Alignment to Access Control: A Unified View of GenAI Policy Enforcement」に付随する拡張論文である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative AI (GenAI) applications have flourished enabling users to chat with large language models, and to create agents to act on their behalf for a variety of tasks. The pace of development of capabilities in this field is incredibly fast with security and safety taking a back seat. Unfortunately, the slower pace at which security and safety mechanisms have evolved has led to real incidents. Policy enables the definition of desirable behavior of applications, and for that reason, it is a cornerstone of making systems secure and compliant. Policy however means different things to different practitioners creating confusion and siloed solutions that are not adequate for compliance. This paper takes a tour of the good, the bad and the ugly when it comes to policy enforcement in GenAI applications. We propose a methodology to systematically analyze and dissect existing approaches to define and enforce policy found in the wild. Based on this principled analysis, we provide recommendations and call for action for the community to address. This paper is a companion extension of USENIX Security 2026 Enigma talk titled "From Alignment to Access Control: A Unified View of GenAI Policy Enforcement" by the author Nathalie Baracaldo.

arXiv ID: 2609.26682 / 要約の誤りについて