arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

画像と質問の根拠を確かめて安全な応答を選ぶ

A$^2$Safe: Counterfactual Evidence-Aligned Adaptive Agent Collaboration for Safe and Effective Visual Question Answering

Quanxing Xu, Ling Zhou, Xian Zhong, Jinyu Tian, Xiaohua Huang, Rubing Huang, Chia-Wen Lin

この論文をやさしく読む

ひとことで言うと

画像か質問のどちらかだけで判断せず、組合せが危険かどうかを根拠付きで確認して、回答・見直し・拒否を選ぶ方法です。

何に役立つ?

画像を扱う対話システムで、危険な回答と無害な質問への過剰な拒否を両方評価する際に役立ちます。報告値はベンチマーク上の成績です。

この研究の面白いところ

危険性に関係のない変更では判断を保ち、危険性を左右する小さな変更では判断を切り替えるよう求めます。単に拒否を増やすのではなく、判断が何に依存しているかを扱います。

どこまで分かった?

要旨ではSIUO、MOSSBench、一般VQAのスコアとトークン負担を報告しています。SIUOの95.72はスコアであり、安全な応答の割合95.72%と読み替えることはできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

マルチモーダル大規模言語モデル(MLLM)による視覚的質問応答(VQA)には、安全で有効な応答を生成するだけでなく、危険性を決めるマルチモーダルな証拠に安全性判断を根拠付けることが求められる。近年の安全性アライメント手法は、回答拒否の振る舞いや文脈に応じた危険性の認識を改善している。しかし、個別には無害な画像と質問の組合せから危険性が生じる場合には特に、安全性の判断結果が正しくても、表面的な文章・画像の相関に依存している可能性がある。 この問題に対して、安全で有効なVQAのために、反実仮想に基づいて証拠を整合させる適応的エージェント協調の枠組みA²Safeを提案する。A²Safeは、局所的な視覚観察、文章の意図、モダリティ間の危険性の関係を、Grounded Safety Evidence Boardに整理し、安全性判断の根拠を明示する。反実仮想を用いた安全性証拠の整合では、安全性に無関係な変更に対する不変性を課す一方、危険性を左右する証拠を最小限変更したときには、安全性の状態と応答モードが適切に切り替わることを求める。 得られた証拠の状態は適応的な協調も支える。根拠となる証拠が十分なら直接回答し、証拠に危険性、不確実性、矛盾がある場合には方針への批評と応答の修正を呼び出す。安全性が重要なVQAと一般的なVQAの相補的な評価手順で、A²SafeはSIUOの安全性スコア95.72を達成し、MOSSBenchで無害な質問への拒否率を14.67%に下げ、トークンの追加負担27.8%で一般VQAの平均スコア78.34を維持した。これらの結果は、安全で有効なマルチモーダル質問応答における、反実仮想による証拠整合を用いた適応的協調の有用性を裏付ける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Visual Question Answering (VQA) with Multimodal Large Language Models (MLLMs) requires not only producing safe and effective responses, but also grounding safety decisions in the multimodal evidence that determines risk. Recent safety-alignment methods improve refusal behavior and contextual risk awareness, yet correct safety outcomes may still rely on superficial textual or visual correlations, particularly when risk emerges from interactions between individually benign image and question content. To address this issue, we propose A$^2$Safe, a counterfactual evidence-aligned adaptive agent collaboration framework for safe and effective VQA. A$^2$Safe organizes localized visual observations, textual intent, and cross-modal risk relations through a Grounded Safety Evidence Board, making the basis of safety decisions explicit. Counterfactual safety evidence alignment enforces invariance to safety-irrelevant changes while requiring appropriate safety-state and response-mode transitions when risk-critical evidence is minimally altered. The resulting evidence state further supports adaptive collaboration, enabling direct answering when grounded evidence is sufficient and invoking policy critique and response revision when evidence is risky, uncertain, or conflicting. Under complementary safety-critical and general VQA protocols, A$^2$Safe achieves a 95.72 SIUO safety score, reduces the benign refusal rate on MOSSBench to 14.67%, and maintains an average general VQA score of 78.34 with 27.8% token overhead. These results support counterfactual evidence-aligned adaptive collaboration for safe and effective multimodal question answering.

arXiv ID: 2609.24098 / 要約の誤りについて