カテゴリ階層のつなぎ方を学習して未知物体を検出する
InterHier: Learning Interconnected Hierarchical Semantics for Open-Vocabulary Object Detection
この論文をやさしく読む
ひとことで言うと
物体カテゴリの上位・下位関係を言葉でモデルへ伝える際、その解釈を助ける文脈を学習する方法です。
何に役立つ?
考えられる用途は、学習時に見ていないカテゴリ名も扱う物体検出です。既存の検出モデルへ組み込める仕組みを提案しています。
この研究の面白いところ
カテゴリ同士を固定の表現でつなぐだけでなく、先頭の文脈を画像領域と整合するように学習し、階層全体の解釈を調整します。
どこまで分かった?
要旨には具体的なベンチマーク名、改善幅、計算費用はありません。最先端との競争力という記述は、あらゆる条件で最良であるという主張ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本論文では、オープン語彙物体検出の階層的な意味表現における、固定の手作りの接続表現の限界を調べる。既存手法は、隣接する上位カテゴリと下位カテゴリの間に固定の接続表現を置くことで、基本カテゴリと未見の新カテゴリの意味的関係を構築する。しかし、この固定した接続表現では、意味階層内の関係を最適に捉えられない可能性がある。 この限界に対処するため、相互接続された階層的意味表現InterHierを提案する。これは、階層関係を含むプロンプトの解釈を全体的に導く、先頭に付けた学習可能な文脈を使う。InterHierは主に二段階で動作する。まず、上位・下位カテゴリを統合して先頭に学習可能な文脈を付け、階層を考慮したプロンプトを構築する。次に、視覚領域の埋め込みとテキスト埋め込みが整合するよう、その文脈を最適化する。InterHierは固定の接続表現に頼る手法より一貫して性能を改善し、既存のオープン語彙物体検出モデルに容易に組み込める。オープン語彙物体検出のベンチマーク実験では、最先端手法に対して競争力のある性能を達成することが示された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 掲載先の記載あり
著者による掲載先の記載:IEEE Access, vol. 14, pp. 14709-14721, 2026。出版社での独立確認は未実施です。
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In this paper, we investigate the limitations of fixed, hand-crafted connectors in hierarchical semantic representations for open-vocabulary object detection. Existing methods establish semantic relationships between base categories and unseen novel categories by placing a fixed connector between adjacent super-/sub-categories. However, such fixed connectors may not optimally capture the relationships within a semantic hierarchy. To address this limitation, we propose interconnected hierarchical semantic representations (InterHier), which utilize a prepended learnable context to globally guide the interpretation of prompts containing hierarchical relationships. InterHier operates in two main stages. First, it constructs a hierarchy-aware prompt by integrating super-/sub-categories and prepending a learnable context. Second, it optimizes this learnable context to align visual region embeddings and textual embeddings. InterHier consistently improves performance over methods that rely on fixed connectors and can be seamlessly integrated into existing open-vocabulary object detection models. Experiments on open-vocabulary object detection benchmarks demonstrate that InterHier achieves competitive performance against state-of-the-art methods.
著者のコメント
12 pages, 6 figures. Published in IEEE Access
arXiv ID: 2609.24026 / 要約の誤りについて