arXiv論文メモ
新着一覧
cs.CR / cs.AI · 査読状況未確認

画像と言語を扱う大規模モデルの出自を指紋で調べる

Fingerprinting Multimodal Large Language Models

Chao Huang, Meng Tong, Kejiang Chen

この論文をやさしく読む

ひとことで言うと

画像と言語を扱うモデルの内部や出力から、派生モデルや蒸留の関係を識別する指紋を作ります。

何に役立つ?

モデルの来歴を監査するための技術研究です。内部にアクセスできる場合と、出力しか見られない場合を分けます。

この研究の面白いところ

内部注意分布の低周波成分を使うAttnPrintと、出力の仮説検定を行うDistillTraceを提案します。19構造・154モデル例で評価しています。

どこまで分かった?

要旨は五種類の下流変更や三種類のパラメータ非依存手法での結果を述べますが、誤判定率の具体値はありません。技術的な関係の証拠と権利侵害の確定は別です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

マルチモーダル大規模言語モデル(MLLM)は多様な画像・テキスト推論タスクを可能にする一方、最近の事例は、不正な配備や無許可の蒸留に対して脆弱であることを示している。モデルの出自を調べる既存の方法は、MLLMが共通の言語バックボーンを持つことで識別が妨げられることが多く、蒸留による違反の検出も難しい。この隔たりを埋め、モデルの所有権を保護するため、マルチモーダルモデルのフィンガープリンティングに関する初めての研究を提示する。 自己注意が低域通過フィルターとして働き、その低周波成分が有益な情報を持つという最近の知見に着想を得て、ホワイトボックスで出自を調べるAttnPrintを開発した。具体的には、モダリティ間の注意分布を抽出し、その低周波成分を分離してモデルの指紋とする。さらに、ブラックボックス監査を可能にするため、MLLMの出力に対する仮説検定を用いて、潜在的なモデル権利侵害を識別するDistillTraceを導入する。 19種類のマルチモーダルアーキテクチャに属する154のモデル個体で広範な実験を行った。特にAttnPrintは、5種類の後段の改変手法に対して頑健性を維持しつつ、派生モデルの検出で高い性能を示した。DistillTraceも、モデルパラメータに依存しない3種類の手法の下で、蒸留関係の証拠を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

While multimodal large language models (MLLMs) enable a wide range of image-text reasoning tasks, recent incidents indicate that they are vulnerable to illicit deployment and unauthorized distillation. Existing solutions for model provenance are typically confounded by shared language backbones in MLLMs and struggle to detect violations of distillation. To bridge this gap and safeguard model ownership, we present the first study on multimodal model fingerprinting. Inspired by recent findings that self-attention acts as a low-pass filter and that its low-frequency components are informative, we develop AttnPrint for white-box provenance. Specifically, we extract cross-modal attention distributions and isolate their low-frequency components to serve as model fingerprints. To facilitate black-box auditing, we further introduce DistillTrace, which employs hypothesis testing of MLLM outputs to identify potential model infringement. We conduct extensive experiments on 154 model instances across 19 multimodal architectures. Notably, AttnPrint achieves strong derivative-model detection performance while remaining robust to five downstream modification techniques. DistillTrace also provides evidence of distillation relationships under three parameter-independent techniques.

著者のコメント

10 pages, 3 figures. Accepted to ACM Multimedia 2026 (MM '26) as an oral presentation

arXiv ID: 2609.20457 / 要約の誤りについて