位置合わせ不要の可視光・熱画像から目立つ物体を検出
S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection
この論文をやさしく読む
ひとことで言うと
可視光画像と熱画像の位置がずれていても、両方の情報を段階的に組み合わせて目立つ物体を見つける手法です。
何に役立つ?
事前に厳密な位置合わせをしにくい可視光・熱画像の組を使う物体検出に役立つと考えられる。要旨で確認されているのは公開ベンチマーク上の性能である。
この研究の面白いところ
全体的な意味情報を先に交換し、その後で局所的なサンプリング位置を補正する順序によって、ずれた画像の情報を融合する。
どこまで分かった?
要旨は複数の公開ベンチマークで競争力のある結果を報告するが、具体的な数値や実環境での運用結果は記載していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
位置合わせをしないRGB画像と熱画像の組から顕著な物体を検出するRGB-T顕著物体検出は、費用のかかる事前の位置合わせを必要としない。しかし、空間的なずれは画素ごとの対応を壊し、異なる画像情報を融合する際に特徴を混ぜてしまう。この問題に対し、意味情報から空間情報へと位置を合わせる枠組みS2Aを提案する。まず、全体情報に導かれる階層的融合モジュールGGHFが、画像全体の意味情報を用いて背景の干渉を抑え、各画像内の階層的特徴を改善する。次に、位置合わせを要しないモダリティ間チャネル注意モジュールAFCAが、チャネルごとの相互作用を通じて補完的な意味情報を全体的に交換し、局所的な位置ずれの干渉を抑える。最後に、空間的な変形可能クロス注意モジュールSDCAが適応的なサンプリング位置のずれを予測し、局所的な画像間の空間対応を回復する。 この意味情報から空間情報へ進む方法では、まず画像間の信頼できる意味情報のやりとりを可能にし、その後で局所的な空間補正を行うことで、位置ずれに起因する特徴の混入を減らす。追加の複雑な仕掛けなしに、複数の公開された位置合わせ不要のRGB-Tベンチマークで高い競争力を持つ性能を達成し、位置ずれによる特徴の混入を緩和できることを示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Alignment-free RGB-T salient object detection (RGB-T SOD) aims to identify salient objects from unregistered RGB and thermal image pairs without costly pre-alignment. However, spatial misalignment breaks pixel-wise correspondence and causes feature contamination during cross-modal fusion. To address this issue, we propose S2A, a semantic-to-spatial alignment framework for alignment-free RGB-T SOD. Specifically, a global-guided hierarchical fusion module (GGHF) first exploits global semantic guidance to suppress background interference and refine hierarchical intra-modal features. Subsequently, the alignment-free cross-modal channel attention module (AFCA) globally exchanges complementary semantic information through channel-wise interaction, effectively overcoming the interference caused by local spatial misalignments. Finally, a spatial deformable cross-attention module (SDCA) predicts adaptive sampling offsets to recover local cross-modal spatial correspondence. Through this semantic-to-spatial paradigm, S2A first enables reliable cross-modal semantic interaction and subsequently performs local spatial calibration, effectively reducing misalignment-induced feature contamination. Without bells and whistles, S2A achieves highly competitive performance on multiple public alignment-free RGB-T benchmarks, demonstrating its effectiveness in alleviating misalignment-induced feature contamination.
arXiv ID: 2609.27413 / 要約の誤りについて