arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

EUのAIリスク分類を評価する公開基盤と画面

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

Jacob T. Emmerson, Phuong-Anh Nguyen-Le, Ronan Romano, Wilber Sean V. Anterola, Yann Billeter, Zhijing Jin

この論文をやさしく読む

ひとことで言うと

EUのAIリスク分類に沿って19の公開ベンチマークを整理し、評価の根拠や集計方法を見比べられる画面を作った研究です。

何に役立つ?

AIモデルのリスク評価で、平均値だけでなく最悪値も確認し、各点数の元になったベンチマークをたどる用途が考えられます。

この研究の面白いところ

18モデルで集計方法を変えると点数が14~37点変わり、採点の一致度や入力変形の妥当性、21人の利用者の反応も調べています。

どこまで分かった?

要旨の数値は対象の18モデル、19ベンチマーク、監査した変形、21人の調査に基づきます。点数だけで実際の被害発生率を測定したとは述べていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

AIの安全性に関する主張はAI分野の外にも広く届くが、裏付けとなる証拠は、入手可能な場合でも不透明だったり、固定された評価に依存したりすることが多い。本論文は、実証的な根拠を一般の人にとってより透明で追跡可能にするための公開評価パイプラインと画面、Systemic Risk Indexを提示する。公開されている19のベンチマークを、EUの汎用AI行動規範が定める四つのシステミックリスク分野、すなわちCBRN、サイバー攻撃、有害な操作、制御喪失に整理する。害の内容を保つような入力の変形と、模擬的な運用状況を用いてモデルを評価する。画面では、平均値と最悪値による集計を切り替え、モデルの能力が総合点に与える影響の設定を変え、各リスク評価をベンチマークの根拠までたどれる。 18モデルを評価すると、最悪値による集計では平均値による集計よりスコアが14~37点低くなり、モデルのリスクを平均で評価すると隠れ得る情報が明らかになった。大規模言語モデルによる判定者と人間の採点者の一致度は、人間同士の一致度に匹敵した(κ=0.78~0.82)。盲検監査では、標本として調べた変形の83%が元の害の内容を保っていた。21人を対象とした調査では、参加者の多くが点数を理解しやすいと回答し、画面によって異なる設定でモデル評価を見るよう促されたと報告した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($\kappa = 0.78\text{--}0.82$), and a blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings

著者のコメント

6 pages, 5 figures, submitted to EACL 2027 Systems Demonstration track

arXiv ID: 2609.28335 / 要約の誤りについて