衛星・航空画像で向きのある物体を文章から特定するモデル群
A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing
この論文をやさしく読む
ひとことで言うと
画像中のどの物体を文章が指すかを、物体の向きも含む枠で特定する手法とデータセットを提案した。
何に役立つ?
衛星・航空画像で斜めに写った物体を文章から探す課題の学習・評価に使えると考えられる。
この研究の面白いところ
文章を使わず候補領域を先に抽出して再利用する方式と、その候補を使う生成型モデルを同じ枠組みに含めた。
どこまで分かった?
要旨は複数のベンチマークで優れた性能と述べるが、比較対象や数値、どの物体で改善したかは記載していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
リモートセンシング画像における視覚的な位置特定は、参照表現で説明された物体を探すことを目的とする。既存手法の多くは水平な境界ボックスを予測するが、さまざまな向きの物体では不正確になりやすい。この問題に対し、互いに補完的な3つの設計からなる、向き付き物体の視覚的位置特定モデル群O²-VGを導入する。 O²-VG-Transは、異なるモダリティをまたぐトランスフォーマーであり、このモデル群の識別型の基盤となる。それを基にしたO²-VG-Uniは、特定の文章による指示なしに、前景にあり得る物体の汎用的な向き付き候補領域を予測する。候補領域の埋め込みを保存しておくことで、物体の検索にも対応する。O²-VG-VLMは、その汎用候補領域を入力の指示として使う自己回帰型の視覚言語モデルであり、複数トークンを予測する方法で、向き付き境界ボックスのトークン群を並列に生成する。 加えて、リモートセンシング画像で向き付き物体を文章から特定するためのデータセットDIOR-R-RSVGを構築する。画像、参照表現、向き付き境界ボックスの組を、学習と評価のために提供する。O²-VGモデル群は、識別型トランスフォーマーから生成型視覚言語モデルまでを含む柔軟な枠組みを提供し、複数のベンチマークで優れた性能を達成した。コードは論文が示す公開リポジトリで提供される。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Visual grounding in remote sensing images aims to locate objects described by referring expressions. Most existing methods predict horizontal bounding boxes, which are often inaccurate for objects with arbitrary orientations. To address this limitation, we introduce O$^2$-VG, a family of models for oriented object visual grounding with three complementary designs. Specifically, O$^2$-VG-Trans is a cross-modality transformer for oriented object visual grounding. It establishes a strong discriminative foundation for the model family. Building upon it, O$^2$-VG-Uni predicts universal oriented proposals for possible foreground objects without specific text prompts. It also supports object retrieval through cached proposal embeddings. Using these universal oriented proposals as input prompts, O$^2$-VG-VLM is an autoregressive vision-language model. It generates oriented box token blocks in parallel through multi-token prediction. In addition, we construct DIOR-R-RSVG, a dataset for oriented object visual grounding in remote sensing images. It provides image, expression, and oriented box triplets for training and evaluation. Together, the O$^2$-VG family provides a flexible framework that spans discriminative transformers and generative vision-language models. It achieves superior performance across multiple benchmarks. Code is available at https://github.com/wokaikaixinxin/ai4rs.
arXiv ID: 2609.28230 / 要約の誤りについて