arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

意味が近い画素をまとめて高スペクトル画像を分類

Token Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification

Yimin Zhu, Mahmood Elahi, Lincoln Linlin Xu

この論文をやさしく読む

ひとことで言うと

多数の波長で撮った画像の画素を、意味が近い部分ごとにまとめて分類するAI手法。

何に役立つ?

高スペクトル画像の細かな領域分類で、空間と波長の情報を効率的に扱う方法として参考になる。

この研究の面白いところ

近くの画素だけで系列を作らず、意味が近い少数のトークンを集め、空間と波長の依存関係を別々に捉える。

どこまで分かった?

性能比較は三つの大規模基準データセットでの実験に基づき、要旨には具体的な数値の差は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

高スペクトル画像(HSI)は豊かな波長と空間の情報を持つが、スペクトルと空間の性質が場所により異なり、空間構造も複雑なため、画素単位で正確に分類するのは難しい。既存の視覚用状態空間モデル(Mamba)は通常、あらかじめ決めた空間的な近傍から系列を作り、意味的な類似性や、場所によって性質が変わることを明示的には考慮しない。そこで、疎なトークンを意味的にまとまった系列へ並べるToken Clustering and Semantic Sequence Mamba(STMamba)を提案する。第一に、大きな構造では、階層的な符号化器・復号器がToken Clustering Module(TCM)で意味的なトークンを段階的に選び、学習するパラメータを持たないCross-scale Neighborhood Attention(CNA)アップサンプラーで密な特徴へ戻す。第二に、細かな構造では、TCMが密度を考慮したクラスタリングで代表的な中心を見つけ、特徴の類似性から各クラスタへの柔らかい所属度を推定する。次に四分木に基づく動的な選択法で、各意味クラスタから空間的に分散した少数のトークンを残す。これにより、一貫した意味的なトークン系列を作りつつ、画素ごとの冗長な表現を減らす。第三に、並列のSpatialとSpectral Semantic-wise Sequencing Mamba(SWSM)モジュールが、同じ意味を持つトークン系列内で、長距離の空間的・波長的な依存関係を相補的に捉え、異なる領域間の無関係な相互作用を抑える。大規模な基準データセット三つでの実験では、定量的・定性的な結果の両面で、STMambaは既存の最先端手法を上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Although hyperspectral images (HSIs) provide rich spectral-spatial information, accurate pixel-level classification remains challenging because of spectral-spatial heterogeneity and complex spatial structures. Existing vision state space models (Mamba) typically construct sequences according to predefined spatial neighborhoods, without explicitly accounting for semantic similarity or spatial non-stationarity. To address this limitation, we propose Token Clustering and Semantic Sequence Mamba (STMamba), which organizes sparse tokens into semantically coherent sequences for hyperspectral image classification with the following features. First, at the macro level, a hierarchical encoder decoder progressively selects semantic tokens with the Token Clustering Module (TCM) and restores dense features using a parameter-free Cross-scale Neighborhood Attention (CNA) Upsampler. Second, at the micro level, TCM first identifies representative cluster centers through density-aware clustering and estimates soft memberships based on feature similarity. A quadtree-based dynamic selection strategy then retains sparse and spatially distributed tokens from each semantic cluster, forming coherent semantic-token sequences while reducing redundant pixel-wise representations. Third, parallel Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture complementary long-range spatial and spectral dependencies within homogeneous semantic token sequences while suppressing irrelevant interactions across heterogeneous regions. Experimental results on three large-scale benchmark datasets demonstrate that STMamba outperforms the SOTA methods with respect to quantitative and qualitative results.

arXiv ID: 2609.28580 / 要約の誤りについて