arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

結果に応じて次の弱点探しを変えるAI安全性評価

CART: Closed-Loop Adaptive Red Teaming for Large Language Models

Dongdong Zhang, Tengchao Lv, Yilin Jia, Yuzhong Zhao, Yupan Huang, Wenshan Wu, Xiangyang Zhou, Shaohan Huang, Nan Yang, Li Dong, Lei Cui, Furu Wei

この論文をやさしく読む

ひとことで言うと

AIの安全性試験で、見つかった失敗に合わせて次の試験内容を変える評価方法です。

何に役立つ?

言語モデルやツールを使うエージェントの弱点を継続して探し、証拠を記録する用途に役立つ。

この研究の面白いところ

試験の作成、試験対象、採点を別の役割とし、それぞれの選択が発見にどう影響するか調べる。

どこまで分かった?

見つかった失敗の数は試験方策の発見能力を示すもので、実運用での失敗頻度を示すものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自動化したレッドチーム評価では、固定した質問を繰り返すことが多い。これは既知の危険を測れるが、検査中に見つけた失敗から学べない。著者らは、各結果を次に何を試すかの判断に使う枠組みCARTを提示する。広い範囲のリスクから始め、現れた弱点を追い、追加の試験に多様性を保ち、各発見の証拠と出所を記録する。試験を作るChallenger、試験されるTarget、結果を評価するJudgeを分離し、それぞれを独立して調べられるようにする。Targetには文字だけのモデルと、利用できるツールが限定されたエージェントを含む。Frontier、JAH、Agenticという三つの評価群で、基準手法が使えるすべてのTargetについて、CARTは固定した初期質問の再実行より多くの失敗と高い平均リスクを発見した。利得はツールを介するエージェントの試験にも及び、文脈に応じた適応が、直接的な質問の反復では試せない弱点を明らかにし得ることを示唆する。ただし、この結果が示すのは試験方策が何を発見したかであり、実運用で失敗がどれほど起こるかではない。また、ChallengerとJudgeの選び方が得られる証拠に影響するため、役割の分離と独立した審査が重要である。CARTは一度限りのチェック項目を、継続的で適応的、監査可能なモデルやエージェントの弱点探索へ変える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Automated red teaming often replays a fixed set of prompts, which measures known risks but cannot learn from failures found during testing. We present CART (Closed-Loop Adaptive Red Teaming), a framework that uses each result to guide what it tests next. CART begins with broad risk coverage, follows weaknesses that emerge, keeps new probes diverse, and records the evidence and source of every finding. It separates the Challenger that creates tests, the Target being tested, which may be a text-only model or a bounded tool-using agent, and the Judge that evaluates the results, allowing these roles to be studied independently. Across three evaluation families (Frontier, JAH, and Agentic), CART discovers more failures and higher average risk than static seed replay for every Target with an available baseline. The gains extend to tool-mediated agent tests, suggesting that contextual adaptation can reveal weaknesses that direct prompt replay does not exercise. These results describe what the test policies discover, not how often failures occur in real deployments. We also find that Challenger-Judge choices affect the evidence uncovered, highlighting the need for role separation and independent review. Overall, CART turns red teaming from a one-time checklist into a continuous, adaptive, and auditable search for model and agent weaknesses.

arXiv ID: 2609.27336 / 要約の誤りについて