衛星画像の説明・質問応答・位置特定を組み合わせて評価
GeoNLI - A Natural Language Interpreter for Satellite Imagery
この論文をやさしく読む
ひとことで言うと
衛星画像を説明し、質問に答え、対象物の位置を示す複数のモデルを組み合わせた研究。
何に役立つ?
衛星画像を使う説明文作成や対象物の位置特定システムを設計するとき、モデルの組合せと指標別の性能を参考にできる。
この研究の面白いところ
説明文作成とVQAではEarthMind、位置特定では複数モデルの多数決を用い、質問の種類ごとの正確度も示した。
どこまで分かった?
評価はVRS BenchとNWPU-VHR-10に基づく。数値質問の正確度52.04%と位置特定64.94%は、用途を判断する際に考慮が必要である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
複数の情報形式を扱い、複数の仕事をこなすモデルは、リモートセンシングのデータセットで高い性能を示している。しかし、学習データが異質で、仕事によってモデルの振る舞いも変わるため、画像の説明文作成、視覚質問応答(VQA)、画像内の対象位置の特定を一つの枠組みで高性能に行うのは難しい。本研究はVRS BenchとNWPU-VHR-10のデータセットで複数のモデルを評価する。EarthMindは説明文作成とVQAの両方で高い結果を示した。対象位置の特定ではRemoteSAM-SAM-v1、RemoteSAM-SAM-v2、DiffuSAMという複数の処理手順を提案し、最終的にはEarthMind、RemoteSAM、SAM3、Falcon、RemoteSAM-SAM3-v1、RemoteSAM-SAM3-v2、DiffuSAMの予測を多数決で組み合わせる。 高度なSAM系モデルとマルチモーダルLLMを組み合わせた、統一的で部品を交換できる処理手順によって、説明文作成、VQA、対象位置特定を実行する。説明文作成の正確度は82%、VQAは83.32%で、質問の種類別には二択90.94%、数値52.04%、意味理解92.06%だった。位置特定の正確度は64.94%だった。多様な視覚言語モデルと独自のRemoteSAM-SAM3モデルを多数決で組み合わせることで、個々の仕事に特化した手法より正確で一貫したリモートセンシング画像の理解が得られる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Multi-modal multitasking models have shown strong performance on remote sensing datasets. However, because these models are trained on heterogeneous data and vary across tasks, designing a unified model that performs well in captioning, visual question answering (VQA), and visual grounding remains challenging. In this work, we evaluate several models on the VRS Bench and NWPU-VHR-10 datasets. The EarthMind model demonstrates strong results in both captioning and VQA. For grounding, we propose multiple pipelines - RemoteSAM-SAM-v1, RemoteSAM-SAM-v2, and DiffuSAM - and ultimately adopt a majority-voting ensemble across EarthMind, RemoteSAM, SAM3, Falcon, RemoteSAM-SAM3-v1, RemoteSAM-SAM3-v2, and DiffuSAM predictions. Our unified, modular pipeline integrates advanced SAM variants with multimodal LLMs to jointly perform captioning, VQA, and grounding. It achieves 82% accuracy on captioning and 83.32% on VQA, with 90.94%, 52.04%, and 92.06% for binary, numeric, and semantic question types respectively. For grounding, it attains 64.94% accuracy. By combining diverse VLMs with our custom RemoteSAM-SAM3 models through ensemble majority voting, the system delivers more accurate and consistent remote-sensing understanding than task-specific approaches.
arXiv ID: 2609.28741 / 要約の誤りについて