港湾動画から少数の船舶検出用画像を選ぶ
Frame-to-Panorama Localization and Context-Aware Sampling for Scene-Specific Ship Detection in a Smart Marina Testbed
この論文をやさしく読む
ひとことで言うと
港の監視動画から船舶検出に役立つ少数の画像を選び、注釈作業を減らす方法。
何に役立つ?
現場のカメラ視野や天候の違いを残しながら、船舶検出器の学習データを効率的に作る助けになる。
この研究の面白いところ
4万超の候補フレームを220画像へ絞り、系列を分けた交差検証で検出性能を測った。
どこまで分かった?
結果は特定のSmart Marinaの動画と検出器での評価。別の港湾やカメラへの転移は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
スマートな海洋インフラからは種類の異なるセンサーの連続データを得られ、実験の繰り返し、デジタルツインの開発、AIを使う海洋サービスが可能になる。ただし現場ごとのモデルを開発するには、センサーだけでは不十分で、過去の動画を空間的に位置付け、状況情報を付け、注釈に有用な少数の画像へ絞る必要がある。本論文は、パン・チルト・ズーム(PTZ)の信頼できる記録がない過去の港湾動画から船を検出するため、各フレームをパノラマへ位置付け、状況を考慮して標本を選ぶ一連の手順を示す。主な貢献は、過去のPTZ動画からカメラの視野情報を復元し、環境情報と映像の多様性を組み合わせ、現場に特化した小さな学習集合を作る、端から端までのデータ整理方法である。具体的にはSuperPointとLightGlueを使って基準パノラマ上にフレームを位置付け、天候と太陽の状態の情報を加え、カメラ視野や環境条件の違いを残す多様性サンプリングで選ぶ。第2段階では、画像を小領域に分けた視覚埋め込みとGaussian混合モデルのクラスタリングにより、水平線付近の例が少ない遠方の船を対象にする。CMMI MDigi-I Smart Marinaで適用すると、注釈候補40,718フレームを220画像へ減らし、99.5%削減した。この集合で追加学習したYOLO26-m検出器は、系列ごとに分けた5分割交差検証で、平均AP50が94.78%±0.51%、平均AP50-95が75.10%±1.73%だった。結果は、重複の多いインフラ動画から、空間と状況に多様性がある小さな学習集合を作り、注釈の労力を大幅に減らしながら現場ごとの検出器へ適応できることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Smart maritime infrastructures provide continuous access to heterogeneous sensing streams, enabling repeated experimentation, digital-twin development, and AI-based maritime services. However, sensing hardware alone is not sufficient for scene-specific model development: historical video streams must also be spatially indexed, contextualized, and reduced to informative subsets for annotation. This paper presents a frame-to-panorama localization and context-aware sampling pipeline for ship detection in historical PTZ maritime video lacking reliable pan, tilt, and zoom metadata. The main contribution is an end-to-end data-curation approach that recovers camera-view information from historical PTZ video and combines it with environmental context and visual diversity to construct compact, scene-specific training sets. Specifically, frames are localized on a reference panorama using SuperPoint and LightGlue, enriched with weather and solar-state metadata, and selected through diversity sampling to preserve variation across camera view and environmental conditions. A second context-aware stage targets under-represented distant-vessel cases near the horizon using tile-level visual embeddings and Gaussian Mixture Model clustering. Applied within the CMMI MDigi-I Smart Marina testbed, the proposed pipeline reduces 40,718 candidate frames to 220 images for annotation, corresponding to a 99.5% reduction. A YOLO26-m detector fine-tuned on this subset achieves a mean AP50 of 94.78% $\pm$ 0.51% and a mean AP50-95 of 75.10% $\pm$ 1.73% under sequence-grouped five-fold cross-validation. These results demonstrate that highly redundant infrastructure video streams can be transformed into compact, spatially and contextually diverse training sets for scene-specific detector adaptation while substantially reducing annotation effort.
arXiv ID: 2609.29447 / 要約の誤りについて