arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

映像の動きで宇宙機の向きの取り違えを減らす

HAT: Hypothesis-Anchored Tracking for Video Monocular Spacecraft Pose Estimation

André Lopo, Atabak Dehban, Rodrigo Ventura

この論文をやさしく読む

ひとことで言うと

一枚では向きを見分けにくい宇宙機について、映像の動きと複数の姿勢候補を組み合わせて追跡する方法です。

何に役立つ?

考えられる用途は軌道上の保守や宇宙ごみ除去での姿勢認識です。報告された検証はデータセット上の評価です。

この研究の面白いところ

誤った候補を早く一つに絞らず、向きの履歴を保持します。過去の出力を書き換えずに逐次推定する設計です。

どこまで分かった?

校正済み画像、実寸CADモデル、対象領域が必要です。数値は4データセット別比較の算術平均で、すべての場面で同じ改善率になることを示すものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

非協力的な対象の単眼6自由度姿勢推定は、軌道上サービスや宇宙ごみ除去に重要である。単一画像による推定器は、ほぼ対称な宇宙機の向きを取り違えることがあり、追跡処理は誤った姿勢を維持してしまう可能性がある。本研究はHypothesis-Anchored Tracking(HAT)を提示する。これは位置合わせと融合の前に、フレーム間の動きを使ってCADに基づく競合する姿勢仮説を選ぶ因果的な枠組みである。各画像で独立に最高スコアの仮説を選ぶ代わりに、競合する向きの履歴を保持し、単眼SLAMが推定した相対軌道の基準となる姿勢を選ぶ。疎な基準点と姿勢融合により、初期化後は過去の出力を修正せず各フレームの推定を行う。 必要なのは、校正済みRGB画像列、実寸スケールを持つCADモデル、検出またはセグメンテーションで与えられる対象画像領域だけである。事前学習済みの姿勢推定ネットワークとSLAMネットワークには、対象専用の学習や微調整を必要としない。MegaPoseとPicoPoseを用いるMega-HATとPico-HATの二方式を、SPARK-2024、SwissCube、SHIRTで評価し、YCB-Videoでは宇宙分野以外の性能を調べる。各手法で一つの時間設定を用いた4データセット別比較の算術平均では、Mega-HATは独立に使うMegaPoseに対して平均姿勢誤差が9.4%低く、持続的に処理できる入力FPSが3.76倍だった。Pico-HATは独立に使うPicoPoseに対して平均姿勢誤差が23.9%低く、FPSが2.42倍だった。SPARK上でのMega-HATのアブレーションとオフライン参照方式により、各構成要素の寄与と、過去の推定を修正することの影響を調べる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Monocular 6-DoF pose estimation of non-cooperative targets is important for on-orbit servicing and debris removal. A single-image estimator can confuse near-symmetric spacecraft orientations, and tracking can preserve an incorrect pose. We present Hypothesis-Anchored Tracking (HAT), a causal framework that uses inter-frame motion to select among competing CAD-based pose hypotheses before alignment and fusion. Rather than independently choosing the highest-scoring hypothesis in each image, HAT retains competing orientation histories and selects a pose to anchor the relative trajectory estimated by monocular SLAM. Sparse anchors and pose fusion provide per-frame estimates after initialization without revising past outputs. The method requires only a calibrated RGB sequence, a metric CAD model, and target image regions, which can be supplied by detection or segmentation. The pretrained pose and SLAM networks require no target-specific training or fine-tuning. We evaluate two versions, Mega-HAT and Pico-HAT, using MegaPose and PicoPose, on SPARK-2024, SwissCube and SHIRT, with YCB-Video assessing performance outside the space domain. Using one temporal configuration per method, the arithmetic means of the four dataset-wise comparisons show 9.4% lower mean pose error and 3.76 times the sustained input FPS for Mega-HAT relative to independent MegaPose, and 23.9% lower mean pose error and 2.42 times the FPS for Pico-HAT relative to independent PicoPose. Mega-HAT ablations on SPARK and an offline reference examine component contributions and the effect of revising past estimates.

著者のコメント

8 pages, 3 figures, 4 tables

arXiv ID: 2609.21597 / 要約の誤りについて