arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

衛星画像モデルを実際の観測データの劣化で評価

RSPDBench: Benchmarking Vision Foundation Models on Earth Observation Tasks Under Physically Grounded Remote-Sensing Product Degradations

Tanjim Bin Faruk, Khondaker Masfiq Reza, Shrideep Pallickara, Sangmi Lee Pallickara

この論文をやさしく読む

ひとことで言うと

衛星などの観測画像に実際の製品処理で起こる欠陥を入れ、画像モデルがどれだけ性能を保つか調べる。

何に役立つ?

地球観測モデルを運用する前に、解像度や位置合わせなどの問題への弱さを評価するのに役立つ。

この研究の面白いところ

単独の欠陥では予測できない複合的な性能低下を示し、最大の単独低下をさらに38ポイント超える例を報告した。

どこまで分かった?

五つのデータセットと指定されたモデルの評価結果であり、すべての観測製品への一般化は要旨からは判断できない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

地球観測向けの画像基盤モデルはきれいな下流ベンチマークで評価されがちだが、運用時の観測データにはモデルへ届く前から、空間解像度、放射量、位置合わせ、ノイズ、データ間の調和に問題が含まれ得る。従来の頑健性評価で使う一般的な画像の破損や広いドメイン変化では、こうした観測製品固有の故障を切り分けられない。RSPDBenchは、物理的な根拠を持つリモートセンシング製品の劣化を使い、画像基盤モデルを評価するベンチマークである。五つの地球観測データセット、七つの基盤モデルの設定、二つの教師ありベースラインについて、監査した個別の劣化と、それらを組み合わせた製品処理の連鎖を評価する。各モデルは、きれいなデータで選ばれた元の評価手順を使い、自身のきれいなデータでの成績からの低下幅で頑健性を測る。劣化への感度には明確な構造があり、解像度を条件にする符号器とチャネルをグループ化する符号器は異なる種類の故障を防ぐ。同じ物理的欠陥が一方のモデルを悪化させ、別のモデルを改善することもある。複数の劣化を組み合わせると、個別の劣化からは予測できない故障が生じ、モデルによって増幅、飽和、特定成分の支配が見られた。最大の単独劣化よりさらに38パーセントポイント成績が下がる例もあった。地球観測モデルの頑健性は、きれいなデータでの精度や一般的な摂動試験だけでは捉えられず、実際の観測製品に含まれる構造的な欠陥でも測る必要がある。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision foundation models targeting Earth observation (EO) tasks are commonly evaluated on clean downstream benchmarks, but operational EO products can already contain spatial, radiometric, alignment, noise, and harmonization defects before reaching the model. Existing robustness evaluations often use generic image corruptions or broad domain shifts, which do not isolate these product-level failure modes. We introduce \textbf{RSPDBench}, a physically grounded \textbf{r}emote-\textbf{s}ensing-\textbf{p}roduct \textbf{d}egradation \textbf{b}enchmark for vision foundation models. RSPDBench evaluates five EO datasets, seven foundation-model entries, and two supervised baselines under audited primitive degradations and compound product chains. Each model is evaluated under its clean-selected native protocol, with robustness measured as the drop from its own clean baseline. Our analysis reveals that degradation sensitivity is strongly structured: resolution-conditioned and channel-grouped encoders protect different failure axes, and the same physical defect can hurt one model while helping another. Compound chains expose failures that isolated degradations do not predict, with model-dependent amplification, saturation, or component dominance, and excess drops up to $38$ percentage points beyond the strongest component. These results show that EO robustness cannot be characterized by clean accuracy or generic perturbation tests alone; it must also be measured against the structured defects that remote-sensing products carry into deployment.

著者のコメント

Accepted to WACV 2027 (Round 1)

arXiv ID: 2609.23427 / 要約の誤りについて