arXiv論文メモ
新着一覧
cs.CR / cs.AI · 査読状況未確認

画像を返すマルチモーダルRAGからのデータ抽出を検証する

Walking the Embedding Space: Datastore Extraction from Multimodal RAG

Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen

この論文をやさしく読む

ひとことで言うと

検索した画像そのものを返すRAGで、繰り返しの画像入力によって内部の画像をどこまで回収できるかを検証します。

何に役立つ?

画像を扱うRAGの情報漏えいリスクを評価し、画像データに対応した保護策を検討するための研究です。

この研究の面白いところ

回収済みの画像を次の入力に利用し、新しい検索結果が出る領域へ適応的に探索を向ける点が特徴です。

どこまで分かった?

対象は画像そのものを応答する構成であり、全種類のRAGに同じ数値が当てはまるわけではありません。回収数は局所特徴の対応に基づく評価です。手法名は要旨で未展開のLaTeX記号になっているため、その表記を保持しました。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

マルチモーダル検索拡張生成(MRAG)は、マルチモーダル大規模言語モデル(MLLM)の生成能力を、関連性があり最新の外部知識に根拠付ける、信頼性が高く費用効率のよい技術として登場した。ハルシネーションの低減など複数の利点がある一方で、個人情報の漏えいやデータ抽出攻撃への脆弱性など、新たな攻撃面も生む。 本論文では、検索された視覚資料そのものが応答となる「画像を返す」MRAGに対して、ブラックボックス環境で動作する適応的かつ自動的なデータ抽出攻撃手順、原文表記\immragを導入する。各クエリは、攻撃者が保有するシャドー画像と、すでにシステムから回収した画像を混ぜ合わせる。関連性で重み付けした再サンプリングにより、後続のクエリを、まだ新しい検索結果が得られる埋め込み空間の領域へ誘導する。テキストのプロンプトに悪意のあるクエリを置き、モデルを情報漏えいへ説得しようとする現在の抽出攻撃と異なり、本手順はユーザーが与える入力画像の中に悪意のある指示を埋め込む。 医療アシスタント、文書中心の支援ツール、汎用ツールという、現実に想定される異なる三つのシナリオで本手順を評価する。実験では、複数のCLIP系検索器に対する攻撃の有効性に加え、さまざまな生成器の影響も調べる。局所特徴の対応に基づく評価では、2,500クエリからなる1回の実行で、最大611枚の異なる放射線画像、566枚の文書スキャン、416枚の汎用画像を再構成し、回収する異なるデータストア項目数は非適応的な基準手法の最大5.6倍に達する。この結果は、マルチモーダルデータ専用の保護策が急務であることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge. Despite presenting several benefits, such as reducing hallucinatory behavior, they also introduce new attack surfaces, including leakage of private information and vulnerabilities against data extraction attacks. In this paper, we introduce $\immrag$, an adaptive and automatic data extraction attack procedure operating in a black box setting against \emph{image-returning} MRAG, a configuration in which the retrieved visual artifact is itself the response. Each query blends an attacker-held shadow image with an image already recovered from the system, and relevance-weighted resampling steers subsequent queries towards regions of the embedding space that still yield novel retrievals. Unlike current extraction attacks that aim to persuade the model towards data leakage by placing a malicious query as a textual prompt, $\immrag$ embeds the malicious instructions inside a user-given input image. We evaluate $\immrag$ on three plausible and distinct real-world scenarios: medical assistant, document-focused helper and general purpose tool. The experiments involve the study of the effectiveness of the attack on multiple CLIP-family retrievers, as well as the impact of various generators. A single 2500-query run reconstructs up to 611 distinct radiology images, 566 document scans and 416 general-purpose images under local-feature correspondence, and reaches up to $5.6\times$ as many distinct datastore items as a non-adaptive baseline. Our results show the urgent need for safeguards specifically designed for multimodal data.

arXiv ID: 2610.01871 / 要約の誤りについて