arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

火星の軌道画像から地表表現を学ぶMarsRecon

MarsRecon: Self-Supervised and Multimodal Surface Representations for Mars

Akshay Naik, Marius F. R. Juston, Jay Mahajan

この論文をやさしく読む

ひとことで言うと

火星の高解像度画像を、地質ラベルを大量に用意せずに学習し、画像と説明文や広域画像を対応付ける仕組みです。

何に役立つ?

火星観測画像をテキストから探したり、局所画像を周囲の広い画像と結び付けたりする検索処理に役立ちます。地質分類などへの利用は、今後の評価が必要な用途です。

この研究の面白いところ

画像の再構成による学習に、位置情報と局所・全体の関係を加えています。テキストから画像への検索と逆方向の検索では、報告された上位10件の再現率に大きな違いがあります。

どこまで分かった?

対象はOlympus Monsの観測です。著者ら自身が、画像の切り出しの重なりをさらに統制し、地質学的な後段課題で評価する必要を挙げています。検索の性能だけで地質解釈の正しさを実証したものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

高解像度の軌道画像は火星表面の豊富な記録を提供するが、地質ラベルが少ないため、教師あり表現学習には制約がある。本研究では、Olympus MonsのHiRISE観測から視覚表現とマルチモーダル表現を学ぶ、地理空間情報を考慮した処理系MarsReconを提案する。処理系はNASA Planetary Data Systemのデータ製品を較正し、地理参照情報を持つ有効な画像パッチを抽出し、ラベルのない画像でマスク付きオートエンコーダを学習する。 主要なStage Aモデル系列では、入力解像度を上げ、無効なトークンを除外することで、学習に使わない評価データでの再構成損失が0.1751から0.1342に低下した。次に視覚エンコーダを固定し、その特徴を観測テキスト、座標、画像の局所・全体文脈と対応付ける。現時点で最も高い性能を示す局所情報中心のモデルは、保持したテスト分割で、画像からテキストへのrecall@10が0.3787、テキストから画像へのrecall@10が0.9161、局所画像から全体画像へのrecall@10が0.4350を達成する。 これらの結果は、火星に特化した事前学習と検索の処理系が機能することを示す。一方、埋め込み表現のより広い有用性を評価するには、切り出した画像間の重なりに対する追加の統制と、後段の地質学的課題での評価が必要である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

High-resolution orbital imagery offers a rich record of the Martian surface, but sparse geological labels limit supervised representation learning. We present MarsRecon, a geospatially aware pipeline for learning visual and multimodal representations from HiRISE observations of Olympus Mons. The pipeline calibrates NASA Planetary Data System products, extracts valid georeferenced patches, and trains a masked autoencoder on unlabeled imagery. Increasing input resolution and filtering invalid tokens reduced held-out reconstruction loss from 0.1751 to 0.1342 in the principal Stage A model series. We then freeze the visual encoder and align its features with observation text, coordinates, and local--global image context. The strongest current local-primary model achieves image-to-text recall@10 of 0.3787, text-to-image recall@10 of 0.9161, and local-to-global recall@10 of 0.4350 on the held-out test split. These results establish a working Mars-specific pretraining and retrieval pipeline; further crop-overlap controls and downstream geological evaluations are needed to assess the broader utility of its embeddings.

著者のコメント

27 pages, 22 figures

arXiv ID: 2609.22379 / 要約の誤りについて