文章の雑音と形状の手掛かりを使って異種画像の物体を照合
Incentive Noise and Structural Prior Infusion for Multi-modal Object Re-Identification
この論文をやさしく読む
ひとことで言うと
撮影方式の違う画像で同じ物体を見つける際に、画像と説明文のずれを調整する方法です。文章を常に正しい手掛かりとみなさず、その曖昧さや形状の違いも扱います。
何に役立つ?
異なる種類の画像を照合するとき、文章情報を追加するだけでは解消しない不一致への対処に役立ちます。要旨では3つのベンチマークによる評価を報告しています。
この研究の面白いところ
雑音を単に取り除くのではなく、画像と文章に応じた雑音を意図的に与えて補完を促します。さらに幾何学的な事前情報と疎な融合を組み合わせ、意味と構造の両面を調整します。
どこまで分かった?
要旨には具体的な精度、比較手法の内訳、雑音の強さごとの結果はありません。3つの評価で有効性が報告されていますが、任意の撮像方式や説明文への性能は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
複数モダリティによる物体の再識別(ReID)では、異なる撮像方式の相補的な情報を利用できる。意味表現をさらに豊かにするため、近年は文章による説明も追加のモダリティとして取り入れられている。しかし、最近の視覚言語手法では、説明文を雑音のない確定的な信号として扱うことが多く、撮像方式に合わない語句や意味の曖昧な表現など、文章に内在する雑音を見落としている。また、主流の手法には、高水準の意味を整合させた後にも残る、モダリティ間の細かな構造の違いを調整する明示的な仕組みがない。 これらの課題に対し、Positive-Incentive Noise(πノイズ)と構造化プロンプトによる調整を中心とする新たな枠組みを提案する。第1に、Semantic Cross-Modal Modulatorは、視覚入力と文章入力の両方を条件とする分布から、タスクを考慮したπノイズを抽出する。これで大域的なトークンを摂動させ、意味情報に導かれたモダリティ間の補完を可能にする。第2に、Structure-Aware Prompt Adapterは、学習可能な幾何学的事前情報をプロンプトを通じて注入し、空間的一貫性を高める。第3に、Context-Aware Sparse Fusionモジュールは、構造的な文脈を抽出して適応的な融合を導き、雑音のある局所的な細部から個体識別の特徴を守る。3つのマルチモーダルReIDベンチマークでの実験により、本手法の有効性と頑健性を示す。コードはhttps://github.com/zw-absin/INSPIで公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic representation, text descriptions have recently been incorporated as an additional modality. However, recent vision-language approaches often treat text descriptions as clean, deterministic signals and overlook their inherent noise, including modality-mismatched phrases and semantically ambiguous expressions. Moreover, prevailing methods lack explicit mechanisms to reconcile fine-grained structural discrepancies between modalities, even after high-level semantic alignment. To address these challenges, we propose a novel framework centered on Positive-Incentive Noise ({\pi}-noise) and structured prompt modulation. First, the Semantic Cross-Modal Modulator harnesses task-aware {\pi}-noise, sampled from a distribution conditioned on both visual and text inputs, to perturb global tokens and enable semantics-guided cross-modal compensation. Second, the Structure-Aware Prompt Adapter injects learnable geometric priors via prompts to enhance spatial consistency. Third, the Context-Aware Sparse Fusion module distills structural context to guide adaptive fusion while shielding identity features from noisy local details. Experiments on three multi-modal ReID benchmarks demonstrate the effectiveness and robustness of our approach. The code is available at https://github.com/zw-absin/INSPI.
著者のコメント
Accepted by ECCV 2026. The version of record may differ slightly
arXiv ID: 2609.24539 / 要約の誤りについて